A mistake I keep seeing in subtitle tools is simple but expensive: someone already has an approved script, but the workflow still starts by transcribing the audio again.
I ran into this while working with scripted voiceovers. The script was already reviewed, but the subtitle tool still wanted to guess the words from audio.
That sounds reasonable at first. Most captioning tools are built around speech-to-text. Upload audio, get words, split them into captions, export an SRT or VTT file.
But scripted video is a different problem.
If the script is already approved, ASR should not be the source of truth. It can help with timing evidence. It can help detect mismatches. But it should not quietly rewrite the words the user already signed off on.







