Nine seconds of silence. That's how long a caller waited after asking our AI assistant a simple question like "What's the TB test process?" long enough that most people would hang up, assuming the call had dropped.
We got that down to about 1.5 seconds. This is the story of how a straightforward platform-integration task turned into designing an in-house Voice AI architecture from scratch and the latency problem that became the real engineering challenge.
Quick summary:
What started as a VAPI integration turned into building an in-house Voice AI stack Twilio for telephony, Deepgram for speech-to-text, Cartesia for voice after customization limits made the managed platform too rigid.
The naive pipeline had ~9 seconds of silence before the caller heard anything. The fix wasn't making things faster overall; it was optimizing for time-to-first-audio instead of total response time.






