Everyone evaluating AI avatar platforms focuses on voice quality. The bigger UX killer is almost always latency — and it's a harder problem than picking a good TTS provider.
Where the delay actually comes from:
User speaks/types
→ STT (if voice input)
→ LLM generates response (streaming helps, but first-token latency matters)






