The hardest part of a personalized read-aloud picture book isn't generating a voice that sounds like the parent. Current zero-shot TTS clears that bar with a surprisingly short reference clip. The hard part is that a voice which sounds correct can still read wrong — flat, evenly paced, no lift on the last line of a page — and a five-year-old notices immediately even though they can't say why.

Here's what matters in that gap: sample quality, prosody, pacing. Plus the constraint that shapes the whole design — consent.

What Zero-Shot Cloning Actually Needs

The mental model most people carry is "more audio equals a better clone." That was true in the fine-tuning era. For current zero-shot architectures — those conditioning a neural codec language model on a short reference — the curve flattens fast. Published work in this family (VALL-E, XTTS and descendants) reports usable speaker similarity from clips measured in seconds, not minutes.

What does not flatten out is sensitivity to sample quality. Worst first: