I Built a System to Mechanically Select Voice Model Candidates
I created a system to mechanically select from voice model candidates. It reads probe sentences with 24 candidate voices, automatically measures and scores them based on Whisper match rate, vowel elongation, speech rate, intonation, and jitter, then selects the top candidates for manual review.
The system allows weighting to be adjusted based on role. For narrators, slower speech rates are given positive weight, while for MCs, wider intonation ranges are favored. As a common penalty, Whisper match rate was weighted between 1.0 and 1.2. Candidates that couldn't read the script accurately were eliminated. This seemed reasonable at the time.
It turns out this was fundamentally wrong in terms of metric design.
There Was a Role with Abnormally Low Match Rates






