The problem with "just add more workers"
Most Spark performance issues on Databricks aren't solved by scaling the cluster — they're caused by shuffle and skew, and no amount of extra nodes fixes a badly partitioned join. This post builds a realistic pipeline (order events joined against a small dimension table, aggregated, and written to Delta Lake) from the ground up, and uses it to work through:
How Spark's shuffle actually behaves during a wide transformation
Diagnosing and fixing data skew with salting and adaptive query execution (AQE)
Laying out the resulting Delta table with Z-Ordering so downstream queries skip irrelevant files









