Running a 27 billion-parameter AI model on a smartphone sounds like claiming you fit a grand piano into a backpack. PrismML, a Caltech spinout backed by Khosla Ventures, says it did exactly that, compressing Alibaba’s Qwen 3.6 large language model from roughly 54 GB down to less than 4 GB and running it on an iPhone 17 Pro with all parameters active.

The company emerged from stealth on March 31, 2026, armed with $16.25 million in seed funding led by Khosla Ventures and a technology that could reshape how AI gets deployed on consumer devices.

What PrismML actually built

PrismML’s proprietary compression techniques, showcased through its Bonsai model family, achieve up to 14x smaller memory footprints and 8x faster inference compared to full-precision models. The model gets radically smaller and runs radically faster, while keeping every single parameter active rather than selectively pruning the model’s capabilities.

CEO Babak Hassibi has emphasized that this technology could enhance AI capabilities without requiring the kind of data center investments that have turned AI infrastructure into a multi-hundred-billion-dollar arms race.