Target Speaker Selection for Neural Network Beamforming in Multi-Speaker Scenarios

We propose a speaker selection mechanism (SSM) to enhance the training of a beamforming neural network. Our approach is motivated by the observation that listeners typically orient themselves toward the target speaker at a slight undershot angle. The mechanism enables the neural network to learn which speaker to focus on in multi-speaker scenarios, based on the relative positions of the listener and speakers. Importantly, only audio input is required during inference. We conduct acoustic simulations to evaluate the effectiveness of the SSM, demonstrating its impact on performance. Results show significant increase in speech intelligibility, quality, and distortion metrics, outperforming both the ideal minimum variance distortionless filter and the same neural network model trained without SSM.

Paper

References (15)

Scroll for more · 3 remaining

Similar papers

© 2026 NYSGPT2525 LLC