Speaker diarization looks solved right up until it isn't. Point a decent model at a clean two-person phone call and it'll nail who said what, turn after turn, with almost no effort. That's the demo everyone shows.
Now drop the same model into a dinner party. Four people, a cross-talk moment, somebody laughing mid-sentence, a "yeah, exactly" fired off while the host is still finishing a thought. The clean-call model falls apart—and it falls apart in specific, nameable ways.
This post is that list. Not "diarization is hard, be careful," but the actual taxonomy: the exact cases that break systems, why each one breaks them, and—the part almost nobody writes down—how you'd even know it's failing. Because here's the uncomfortable part: the metric most people quote barely notices some of the worst failures.
Why diarization is still hard
Diarization is two jobs stapled together. First, figure out how many distinct speakers exist. Second, draw the boundaries that say person A talked from here to here, then person B took over. Do both perfectly and you get a clean speaker-labeled transcript. Miss on either and the errors compound.






