Voice Arena and Hugging Face partner to launch open ASR evaluation for Hindi and Indian English
Benchmarks decide what gets built. A model that scores well on the Open ASR Leaderboard gets adopted and iterated on, while capabilities the leaderboard does not measure tend not to improve. Much of the recent work on the leaderboard has gone into making the evaluation metrics more trustworthy:
Held-out private splits.
Benchmark-fitting analysis to quantify how much models are reproducing reference transcripts rather than transcribing solely on the audio.
Closing the gaps in normalisers to ensure correct predictions/variants are not penalized.






