Short answer: the cheapest speech-to-text API for an EU startup is the one that passes the startup's region, quality, latency, and recovery gates, then produces the lowest measured cost for the startup's own audio mix; a public per-minute price cannot establish that result by itself.

Treat this as a procurement runbook, not a rate-card contest. Freeze the workload, collect dated quotes for the exact modes under review, replay one corpus through each candidate, and preserve the raw results. OpenAI, Deepgram, AssemblyAI, and Google Cloud can sit in the candidate column, but no name gets a pass on the same acceptance criteria.

The order matters. A low rate attached to a deployment that fails an EU data boundary, transcription-quality target, or recovery drill is not a bargain. It is an excluded option.

What should an EU startup measure when comparing speech-to-text API per-minute pricing?

Begin with a capacity envelope. Record uploaded audio minutes per month, peak uploads per minute, concurrent live sessions if streaming is in scope, the p50 and p95 recording duration, language mix, channel count, and the share of audio expected to need replay. Don't collapse those inputs into one average recording: ten thousand short voice notes and a few long interviews can consume the same audio minutes while creating very different queue pressure, request volume, and failure recovery work.