IBM released granite-speech-5.0-470m-turboctc, a pair of compact, encoder-only speech recognition models built from a stack of 16 Conformer blocks with self-conditioning at the eighth block and chunkwise attention to avoid quadratic scaling with sequence length. Strided-convolution downsampling reduces 100-frames-per-second spectrograms to a 12.5-tokens-per-second output rate, which is decoded directly via CTC rather than through a separate language model. IBM reports the models were trained on roughly 57,610 hours of natural speech data plus 2,740 hours of synthetic data, and reach over 12,600 RTFx on an Nvidia H200 GPU while remaining competitive with larger models on transcription benchmarks.