When developing a voice conversion app, you often find that the quality of the conversion depends more on the choice of the target voice (anchor) than on the conversion model itself. In our voice conversion app, we used 4 to 12 seconds of recordings from voice actors or narrators as anchors. Initially, we relied on "what sounded good to the ear" to select these segments, but this approach didn’t scale.
Manually listening to and selecting good segments from dozens of recordings was not only tedious but also inconsistent, as judgments varied based on mood. Moreover, the quality of recordings was inconsistent: some had overlapping voices in interviews, excessive emotional intonation, or noise. Continuously filtering out "unsuitable anchor segments" manually was impractical.
This article documents how we replaced the "ear-based selection" with a three-stage filter system. The goal was to mechanically select segments that are "single-speaker, natural, and high-quality." As a result, the minimum PESQ of the anchor set improved from 1.24 to 2.31, and we ultimately compiled over 70 anchors.
Prerequisites: Three Criteria for Anchors
What makes a good anchor? For voice conversion, the target audio must meet the following three conditions simultaneously:






