Smallest.ai closed a $13 million Series A round this week. The capital brings their total funding to $21 million since they launched late in 2024. But the money is the least interesting part of the announcement. The real story is the architectural bet they are making with their new Hydra model.

They are betting that the next leap in voice agents will not come from making large language models faster. It will come from using smaller, specialized models built exclusively for human conversation.

The thesis is simple. While you speak to me, I am already thinking. I might interrupt you if you talk for too long. Humans listen, think, and respond in parallel. Large language models do not. They wait for the entire prompt, and then they start thinking. In a text chat, you tolerate that latency. On a phone call, a 500-millisecond pause feels broken.

Splitting the voice agent in two

Smallest.ai is building an asynchronous architecture they call Voice 4.0. The core idea is to split the agent's brain.