AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation,\n Recognition and Speaker Diarization in Conference Scenario

In this paper, we present AISHELL-4, a sizable real-recorded Mandarin speech\ndataset collected by 8-channel circular microphone array for speech processing\nin conference scenario. The dataset consists of 211 recorded meeting sessions,\neach containing 4 to 8 speakers, with a total length of 120 hours. This dataset\naims to bridge the advanced research on multi-speaker processing and the\npractical application scenario in three aspects. With real recorded meetings,\nAISHELL-4 provides realistic acoustics and rich natural speech characteristics\nin conversation such as short pause, speech overlap, quick speaker turn, noise,\netc. Meanwhile, accurate transcription and speaker voice activity are provided\nfor each meeting in AISHELL-4. This allows the researchers to explore different\naspects in meeting processing, ranging from individual tasks such as speech\nfront-end processing, speech recognition and speaker diarization, to\nmulti-modality modeling and joint optimization of relevant tasks. Given most\nopen source dataset for multi-speaker tasks are in English, AISHELL-4 is the\nonly Mandarin dataset for conversation speech, providing additional value for\ndata diversity in speech community. We also release a PyTorch-based training\nand evaluation framework as baseline system to promote reproducible research in\nthis field.\n

Paper

References (34)

Scroll for more · 22 remaining

Similar papers

© 2026 NYSGPT2525 LLC