A speaker encoder turns a few seconds of speech into a vector, and the distance between two vectors is treated as an answer to "are these the same person." That number gets used for authentication, for voice-cloning pipelines, and increasingly as the yardstick in papers asking whether a synthetic voice preserved someone's identity.

The number is not neutral. It is a reading from an instrument, and I wanted to know how much the instrument was contributing.

So I ran fourteen of them over an identical, frozen list of 11,935 trials.

The corpus is why the question is answerable

Most emotional-speech corpora use different sentences for different emotions. A model can then learn angry vocabulary instead of angry delivery, and you cannot separate them afterwards.