A voice assistant can be technically correct and still feel broken.

The failure often begins after the user finishes speaking. Nothing dramatic happens. There is simply a pause: long enough for the user to wonder whether the system heard them, but short enough for each service dashboard to report an acceptable result.

That pause is the product.

Users do not experience speech recognition, a language model, a tool call, and speech synthesis as separate services. They experience one conversational turn. If any handoff is late, uncertain, or difficult to cancel, the whole interaction begins to feel like a phone tree rather than a conversation.

The useful question is therefore not, "Which model is fastest?" It is: