Every downstream artefact — show notes, chapters, clips, the written version, the search index — is a transformation of the transcript. So transcript accuracy is not one task among several. It is the input to all of them, and an error there is an error everywhere.
Transcription first, everything downstream
The correct order of work is transcript, correction, then everything else, and the correction step is the one people skip because the raw output looks fine when you read a paragraph of it.
Record with separate tracks per speaker if you possibly can.
Transcribe with speaker labels and word-level timings. Word-level rather than segment-level, because chapters and clip extraction both need to find a moment rather than a paragraph.








