Improving Multi-Speaker Transcription for Live News Broadcasts With Canary 1B and Pyannote Diarization

Live news programs are especially challenging for Automatic Transcription (AT) systems as they contain numerous and intricate interactions between presenters, reporters, voiceovers and off-site clips. In this paper, we investigate whether Automatic Speech Recognition (ASR), and instrumental time-aligned labels from ASR decoding paths, can be used to improve speaker diarization in domain-specific data. We utilize NVIDIA's Canary 1B ASR model within the WhisperX framework in order to improve AT performance, word level alignment and speaker segmentation. Canary 1B has generally shown stronger performance based on average WER and RTFx in the previous evaluation. Using experimental dataset from an English broadcast with several speakers, our proposed model achieves 0% Diarization Error Rate (DER), 2.2% Word Error Rate (WER), and 0.04% Character Error Rate (CER). These results demonstrate that by integrating Canary 1 B transcription model with our search-based decoder can provide a more effective and efficient transcription system for live news applications.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC