IBM Says Granite Speech 5.0 Transcribes 3.5 Hours of Speech in One Second

IBM released two compact English speech recognition models on August 25, 2026, claiming transcription throughput no open model has posted before: more than 3.5 hours of speech processed in a single second. The pair, Granite Speech 5.0 TurboCTC and a noncommercial counterpart, each carry just 470 million parameters, and the company reports aggregate throughput above 12,600 RTFx on a single NVIDIA H200 GPU, meaning audio flows through the models more than twelve thousand times faster than real…

This article has been indexed from Unite.AI

Read the original article: