This is a record of designing voices from single-line captions, automatically creating a learning corpus, and passing all 12 role-specific voices (narrator/counselor/sales/presenter/operator/MC for both men and women) through full inspection. I wrote about the failures I encountered during approximately one month of actual work, divided into 18 articles. This article is the table of contents.

The Conclusion Upfront

Voice design, voice manufacturing, and voice operation are different technologies with different failures.

Design uses diffusion TTS. The voice is determined by the caption and random seed, making it fully reproducible.

Manufacturing is primarily about corpus generation. The design of the quality gate directly determines the voice quality.