The Object Navigation (ObjectNav) task requires an agent to locate a specified target in an unseen environment. Without prior knowledge of the layout, the agent must perform semantic reasoning to infer the target’s potential location based on environmental memory accumulated during navigation. Previous studies have indicated that predicting potential locations of target objects based on known maps is crucial for ensuring ObjectNav success and improving efficiency. Diffusion models have demonstrated the ability to learn distributional relationships among features in RGB images, thereby generating novel and realistic images. However, directly training a diffusion model to complete unknown areas from partial semantic maps often leads to poor convergence and limited generalization. In this work, we propose a self-supervised autoencoder designed for indoor semantic maps, which compresses high-dimensional large-scale maps into low-dimensional latent features. These features can be reconstructed back into the original semantic maps via the decoder. We then train a diffusion model in this latent space to perform the map completion task. Finally, we create a dedicated benchmark dataset based on common indoor navigation datasets to evaluate map completion performance, and compare our method with other state-of-the-art approaches to demonstrate its effectiveness.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex