Now available in Databricks Runtime 19 Beta

by Wenchen Fan, Andreas Neumann, Serge Rielau, Szehon Ho, Gengliang Wang, Linhong Liu, Hyukjin Kwon, Jerry Peng, DB Tsai, Xiao Li and Reynold Xin

Apache Spark 4.2 moves more of the modern data and AI stack into the engine itself. Building on Spark 4.x, the release adds governed metrics, vector and top-K primitives, a more Arrow-first Python path, first-class change data capture, and stronger streaming and operational foundations.

This makes Spark more useful on both sides of an AI application. It improves the quality and freshness of the data supplied to AI agents, and it makes Spark easier for applications and agents to invoke as a remote execution service. The AI story is concrete: trusted semantics, native retrieval primitives, fresh change data, and open interfaces to Spark-scale computation.

Spark 4.2 can be understood through four benefits: