NTT Speaker Diarization System for Chime-7: Multi-Domain, Multi-Microphone end-to-end and Vector Clustering Diarization

This paper details our speaker diarization system designed for multi-domain, multi-microphone casual conversations. The proposed diarization pipeline uses weighted prediction error (WPE)based dereverberation as a front end, and separately applies end-to-end neural diarization with vector clustering (EEND-VC) to each channel. It integrates the diarization result obtained from each channel using diarization output voting error reduction plus overlap (DOVER-Lap). To harness the knowledge from the target domain and the results integrated across all channels, we apply self-supervised adaptation for each session by retraining the EEND-VC with pseudo-labels derived from DOVER-Lap. We incorporated our proposed system into NTT’s submission for a distant automatic speech recognition task in the CHiME-7 challenge. Our system obtained third place in the diarization performance by improving the development and evaluation sets by 65 % and 62 % compared to the organizer-provided, VC-based baseline diarization system.

Paper

Similar papers

© 2026 NYSGPT2525 LLC