Until now, running complex ML workloads inside a data clean room meant hitting a wall. Most clean room environments are limited to SQL queries or single-node Python, which runs into memory constraints quickly at enterprise data volumes. Teams ended up treating clean rooms as a compliance tool rather than a place to build models. ML Jobs changes that. ML Jobs in Snowflake Data Clean Rooms™ is now generally available.

Data scientists can now bring their standard Python ML stack with distributed training, hyperparameter optimization, custom packages and GPU compute directly into a multiparty collaboration. Models train on combined data from multiple organizations without raw records leaving anyone's account, and the pipeline runs automatically rather than requiring manual intervention each time.

Consider a concrete example from advertising. An advertiser builds audience and measurement models using ad log data from publishers, identity resolution data from identity providers and transaction signals from retail data partners. Each data source improves model quality. However, each provider has legitimate concerns about how their data is used and who can see it. At the same time, the advertiser's model logic and scoring algorithms are proprietary IP they have no interest in exposing to their data partners.