Thinking Machines Lab has released a research preview of its first AI model, designed to break voice AI out of the traditional question-and-answer pattern. The model processes audio, video, and text in parallel 200-millisecond chunks, and the startup claims it beats OpenAI's GPT-Realtime-2 and Google's Gemini Live on interaction quality.
Thinking Machines Lab has published a research preview of what it calls Interaction Models, AI models that handle interaction natively rather than through external scaffolding. The core idea is that interactivity should scale alongside intelligence, not get treated as an afterthought.
Current voice AI systems still feel robotic
Today's real-time systems like GPT-Realtime or Gemini Live continuously take in audio, but the actual language model never sees it directly. According to Thinking Machines, a "harness" of separate components sits in front of the model, including things like a voice activity detector that decides when a speaker's turn is over. Only then does the finished utterance get handed to the model, which generates a complete response. While it's talking, its perception freezes, receiving no new information until it finishes or gets interrupted.







