Rendering a document as a two-host podcast is not just alternating two TTS voices line by line. Do that and you get something that sounds like two people reading unrelated scripts in the same room.
The problem is prosody, not voices
Human conversation carries information in timing. A reply that arrives instantly reads as agreement. A short gap reads as consideration. Speakers overlap slightly at turn boundaries. Pitch tends to fall at the end of a statement and rise before a handoff.
Concatenated TTS has none of this. Each utterance is synthesised in isolation with neutral prosody and identical gaps, and the result sits in an uncanny valley — clearly speech, clearly not conversation.
Things that measurably help






