The Granite Speech family has expanded with two new models designed to push the boundaries of transcription speed. The Granite Speech 5.0 Turbo CTC models, featuring only 470 million parameters, achieve unprecedented performance, reaching over 12,600 RTFx on an NVIDIA H200 GPU. This allows them to transcribe more than 3.5 hours of speech in a single second using batched inference.

Two Models, Different Licensing

The release consists of two variants that differ primarily in their training data and licensing terms:

  • granite-speech-5.0-470m-turboctc: Trained on a smaller dataset and released under the Apache 2.0 license for commercial use.
  • granite-speech-5.0-470m-turboctc-nc: Trained on additional data and carries a CC-BY-NC-SA-4.0 (non-commercial) license.

Unofficial results from the OpenASR Leaderboard show the non-commercial model achieving a Word Error Rate (WER) of 4.85%, while the Apache 2.0 version scored 5.00%. Both offer high accuracy despite their compact size.

Architecture Focused on Efficiency

Unlike previous Granite Speech models that combined an acoustic encoder with a language model (LM), these new versions are encoder-only. This design provides over 20x faster throughput and a significantly smaller memory footprint, making them ideal for speech-to-text tasks on edge devices.

Technically, the models utilize a stack of 16 Conformer blocks and three stages of 2x subsampling to reduce the token rate to 12.5 tokens per second. Training involved a mix of natural and synthetic datasets, including 2,000 hours of multi-speaker data and specialized sets containing numbers, currencies, and addresses generated via LLMs and synthesized using StyleTTS2.