In the first post I argued there are two ways to pull meaning out of audio:

measure it with signal processing, or estimate it with a model. This post

is the story of a problem where the obvious move was to estimate — and where

measuring turned out to be better.

The problem: labeling who is speaking. A transcript that says "Agent: …" and