Automatic Speech Diarization and Recognition (ASDR) makes ASR system in real-word scenarios more practical by assigning speaker labels to transcribed texts. However, it faces unique challenges due to factors such as speaker overlap, background noise, etc. In this report, we propose a cascaded ASDR system, which fully leverage pretrained feature extraction model. It narrows the cp-CER gap of ASDR system with oracle diarization and our ASDR system to merely 0.33%. Experiment on ICMC challenge shows that our system achieves a cp-CER of 26.37%, which ranked 3rd in the Track II ASDR task.
Paper
Full text
Ximalaya ASDR System for ICASSP 2024 in-Car Multi-Channel (ICMC) ASR Challenge
Semantic Scholar · Computer Science · 2024
Abstract
Automatic Speech Diarization and Recognition (ASDR) makes ASR system in real-word scenarios more practical by assigning speaker labels to transcribed texts. However, it faces unique challenges due to factors such as speaker overlap, background noise, etc. In this report, we propose a cascaded ASDR system, which fully leverage pretrained feature extraction model. It narrows the cp-CER gap of ASDR system with oracle diarization and our ASDR system to merely 0.33%. Experiment on ICMC challenge shows that our system achieves a cp-CER of 26.37%, which ranked 3rd in the Track II ASDR task.