In the first post I argued there are two ways to pull meaning out of audio:
measure it with signal processing, or estimate it with a model. This post
is the story of a problem where the obvious move was to estimate — and where
measuring turned out to be better.
The problem: labeling who is speaking. A transcript that says "Agent: …" and






