For most of the last decade, transcription was something you did after the fact.
You recorded a call, dumped the file into a queue, waited a few minutes (or a few hours), and got back a block of text. That was the deal. Batch, async, post-hoc—whatever you want to call it, the audio was already over by the time the model saw it. And for a long time, that was fine, because that's all the technology could reliably do.
Recently, the field has quietly crossed the line where the most interesting, highest-value speech-to-text work happens while people are still talking. Realtime isn't a niche feature you bolt on for a captions demo. It's becoming the default that everything else gets measured against.
Here's the argument.
The shift: the interesting STT workloads are now live






