TL;DR

Today we're delighted to announce the release of two new models in the Granite Speech family: compact, 470M-parameter English speech recognition models that pair strong accuracy with unprecedented speed — over 12,600 RTFx on an NVIDIA H200 GPU, meaning they can transcribe more than 3.5 hours of speech in one second using batched inference.

granite-speech-5.0-470m-turboctc

granite-speech-5.0-470m-turboctc-nc

To get a sense of the models' responsiveness, check out our WebGPU demo of streaming speech recognition. Note that the demo only runs on Chrome or Edge browsers.