Most transcription APIs still force you to do this:

Download the YouTube video

Extract the audio with yt-dlp + ffmpeg

Upload the file

Call the transcription endpoint