Three-class Overlapped Speech Detection using a Convolutional Recurrent Neural Network

In this work, we propose an overlapped speech detection system trained as a\nthree-class classifier. Unlike conventional systems that perform binary\nclassification as to whether or not a frame contains overlapped speech, the\nproposed approach classifies into three classes: non-speech, single speaker\nspeech, and overlapped speech. By training a network with the more detailed\nlabel definition, the model can learn a better notion on deciding the number of\nspeakers included in a given frame. A convolutional recurrent neural network\narchitecture is explored to benefit from both convolutional layer's capability\nto model local patterns and recurrent layer's ability to model sequential\ninformation. The proposed overlapped speech detection model establishes a\nstate-of-the-art performance with a precision of 0.6648 and a recall of 0.3222\non the DIHARD II evaluation set, showing a 20% increase in recall along with\nhigher precision. In addition, we also introduce a simple approach to utilize\nthe proposed overlapped speech detection model for speaker diarization which\nranked third place in the Track 1 of the DIHARD III challenge.\n

Paper

References (31)

Scroll for more · 19 remaining

Similar papers

© 2026 NYSGPT2525 LLC