Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments

Although deep reinforcement learning (DRL) approaches in audio signal processing have seen substantial progress in recent years, fully audio-driven DRL for tasks such as navigation, gaze control and head-orientation control have received little attention. Yet, such audio-driven DRL approaches are highly relevant for the development of fully autonomous audio-based agents as they can be seamlessly merged with other audio-driven (DRL) approaches such as automatic speech recognition and emotion recognition. Therefore, we propose an end-to-end, audio-driven DRL framework in which we utilize deep Q-learning to develop an autonomous agent that orients towards a talker in the acoustic environment based on stereo speech recordings. Our results show that the agent learned to perform the task in a range of naturalistic acoustic environments with varying degrees of reverberation. Quantifying the degree of generalization of the proposed DRL approach across acoustic environments revealed that policies learned by an agent trained on medium or high reverb environments generalized to low reverb environments, but policies learned by an agent trained on anechoic or low reverb environments did not generalize to medium or high reverb environments. Taken together, this study demonstrates the potential of fully audio-driven DRL for tasks such as head-orientation control. Furthermore, our findings highlight the need for training strategies that enable robust generalization across acoustic environments in order to develop real-world audio-driven DRL applications1.

Paper

References (30)

Scroll for more · 18 remaining

Similar papers

© 2026 NYSGPT2525 LLC