Everyone evaluating AI avatar platforms focuses on voice quality. The bigger UX killer is almost always latency — and it's a harder problem than picking a good TTS provider.

Where the delay actually comes from:

User speaks/types

→ STT (if voice input)

→ LLM generates response (streaming helps, but first-token latency matters)