Alex Smola, the machine learning researcher who co-founded Boson AI in 2023, is preparing to release the company’s first speech-to-speech model called Higgs RealTime. His thesis is straightforward: voice AI models have finally crossed the quality threshold needed to fundamentally change how humans interact with machines.

What Boson AI is building

Boson AI, headquartered in Santa Clara, California, was co-founded by Smola and Mu Li with a focus on creating foundation models for natural, responsive voice interactions.

The company’s Higgs TTS 3, released on June 4, 2026, supports expressive conversational speech in over 100 languages. It features zero-shot voice cloning, meaning it can replicate a voice without needing extensive training samples, and inline emotion control that lets developers dial up or down the emotional register of generated speech in real time.

Alongside the text-to-speech model, Boson AI launched its Higgs Avatar API in June 2026. That tool generates real-time talking-head video from a single still image paired with audio or text inputs.