End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings
We present an end-to-end deep network model that performs meeting diarization\nfrom single-channel audio recordings. End-to-end diarization models have the\nadvantage of handling speaker overlap and enabling straightforward handling of\ndiscriminative training, unlike traditional clustering-based diarization\nmethods. The proposed system is designed to handle meetings with unknown\nnumbers of speakers, using variable-number permutation-invariant cross-entropy\nbased loss functions. We introduce several components that appear to help with\ndiarization performance, including a local convolutional network followed by a\nglobal self-attention module, multi-task transfer learning using a speaker\nidentification component, and a sequential approach where the model is refined\nwith a second stage. These are trained and validated on simulated meeting data\nbased on LibriSpeech and LibriTTS datasets; final evaluations are done using\nLibriCSS, which consists of simulated meetings recorded using real acoustics\nvia loudspeaker playback. The proposed model performs better than previously\nproposed end-to-end diarization models on these data.\n
Paper
References (32)
Scroll for more · 20 remaining