IBM released two compact English speech recognition models on August 25, 2026, claiming transcription throughput no open model has posted before: more than 3.5 hours of speech processed in a single second. The pair, Granite Speech 5.0 TurboCTC and a noncommercial counterpart, each carry just 470 million parameters, and the company reports aggregate throughput above 12,600 RTFx on a single NVIDIA H200 GPU, meaning audio flows through the models more than twelve thousand times faster than real…
Read the original article: