Only the voice-converted audio was inexplicably stretched during playback—the cause was not the model itself, but Whisper's 30-second limit used for semantic extraction. We present the investigation from symptoms to root cause and the solution through chunking.