This technical report describes VibeVoice-ASR-Streaming, an LLM-based end-to-end system for real-time, speaker-attributed speech recognition that processes fixed-size audio chunks plus a small amount of lookahead audio and prior text, unifying transcription and speaker attribution in a single streaming model rather than two separate pipelines. The team evaluated 1.5B and 7B parameter versions; the larger model achieved the lowest average word and character error rates across five benchmarks and led on speaker-attribution accuracy in most evaluation settings.
