A speaker encoder turns a few seconds of speech into a vector, and the distance between two vectors is treated as an answer to "are these the same person." That number gets used for authentication, for voice-cloning pipelines, and increasingly as the yardstick in papers asking whether a synthetic voice preserved someone's identity.
The number is not neutral. It is a reading from an instrument, and I wanted to know how much the instrument was contributing.
So I ran fourteen of them over an identical, frozen list of 11,935 trials.
The corpus is why the question is answerable
Most emotional-speech corpora use different sentences for different emotions. A model can then learn angry vocabulary instead of angry delivery, and you cannot separate them afterwards.






