Bayesian HMM clustering of x-vector sequences (VBx) in speaker\n diarization: theory, implementation and analysis on standard tasks

The recently proposed VBx diarization method uses a Bayesian hidden Markov\nmodel to find speaker clusters in a sequence of x-vectors. In this work we\nperform an extensive comparison of performance of the VBx diarization with\nother approaches in the literature and we show that VBx achieves superior\nperformance on three of the most popular datasets for evaluating diarization:\nCALLHOME, AMI and DIHARDII datasets. Further, we present for the first time the\nderivation and update formulae for the VBx model, focusing on the efficiency\nand simplicity of this model as compared to the previous and more complex BHMM\nmodel working on frame-by-frame standard Cepstral features. Together with this\npublication, we release the recipe for training the x-vector extractors used in\nour experiments on both wide and narrowband data, and the VBx recipes that\nattain state-of-the-art performance on all three datasets. Besides, we point\nout the lack of a standardized evaluation protocol for AMI dataset and we\npropose a new protocol for both Beamformed and Mix-Headset audios based on the\nofficial AMI partitions and transcriptions.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC