Part of qwen3-tts — a pure C inference engine for Qwen3-TTS.
TL;DR
Qwen3-TTS ships 9 neutral preset speakers. No emotion control, and cloning a voice used to mean a huge file that couldn't emote at all. We changed both:
Small clones. A cloned voice is now a ~25 MB .qvoice "graft" — small enough to share, and, crucially, built so the emotion machinery still works on it.
Emotions on any voice. One flag — --emotion <sad|joy|anger|fear|disgust|surprise> — works on presets and cloned voices, in every supported language.







