The quality is legitimately good now — calm NPC dialogue and narration are close to indistinguishable from a real recording. But there's a licensing gotcha worth knowing before you generate a single line: the free tier has no commercial rights. Anything you make on it requires attribution and can't go into a monetized game. Starter at $5/month is the actual floor for shippable audio, not the free tier most people default to first.
Two workflow notes that aren't obvious from the docs:
Voice Design vs cloning — Voice Design generates a synthetic voice from a text description ("gruff middle-aged man, slight Eastern European accent"), no recordings needed, and it's more stable across generations than a clone. Instant cloning needs 1-3 minutes of clean audio and works fine for short lines, but drifts slightly on long or unusual sentences. For a main character with hundreds of lines, that drift matters — worth testing across your actual dialogue range before committing to a voice for production.
Fantasy names are the recurring pain point. Anything outside standard English phonemes gets mispronounced inconsistently between generations. Spell it phonetically in the input text ("Xrathul" → "Zrathool") rather than fighting the model on the literal spelling.







