Knowledge-driven Scene Priors for Semantic Audio-Visual Embodied Navigation

Generalisation to unseen contexts remains a challenge for embodied navigation\nagents. In the context of semantic audio-visual navigation (SAVi) tasks, the\nnotion of generalisation should include both generalising to unseen indoor\nvisual scenes as well as generalising to unheard sounding objects. However,\nprevious SAVi task definitions do not include evaluation conditions on truly\nnovel sounding objects, resorting instead to evaluating agents on unheard sound\nclips of known objects; meanwhile, previous SAVi methods do not include\nexplicit mechanisms for incorporating domain knowledge about object and region\nsemantics. These weaknesses limit the development and assessment of models'\nabilities to generalise their learned experience. In this work, we introduce\nthe use of knowledge-driven scene priors in the semantic audio-visual embodied\nnavigation task: we combine semantic information from our novel knowledge graph\nthat encodes object-region relations, spatial knowledge from dual Graph Encoder\nNetworks, and background knowledge from a series of pre-training tasks -- all\nwithin a reinforcement learning framework for audio-visual navigation. We also\ndefine a new audio-visual navigation sub-task, where agents are evaluated on\nnovel sounding objects, as opposed to unheard clips of known objects. We show\nimprovements over strong baselines in generalisation to unseen regions and\nnovel sounding objects, within the Habitat-Matterport3D simulation environment,\nunder the SoundSpaces task.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC