Synergistic Fusion for Traversability: Combining Footsteps, Semantics, and Geometry

Acquiring dense, pixel-wise traversability labels for quadruped robots is a significant bottleneck that relies heavily on costly and time-consuming manual annotation, limiting the scalability and adaptability of robots. To address this, we propose a novel autolabeling pipeline that generates high-quality traversability labels without human intervention. Our method synergistically fuses three key sources of information available to the robot: (1) weak supervision from the robot’s own future footstep plans, derived from SLAM, to guide a video-native segmentation model; (2) closed-set semantic context from a Mask2Former model to disambiguate surfaces; and (3) geometric cues from an RGB-D sensor to ensure sharp boundary alignment. Through experiments conducted in diverse urban and natural environments in Daejeon, South Korea, we demonstrate that a lightweight segmentation model trained exclusively on our automatically generated labels significantly outperforms strong baselines, including state-of-the-art models pre-trained on large-scale public datasets like Mapillary Vistas. Notably, our approach mitigates domain shift and achieves a 0.11 higher mIoU than a SAM-prompted single-frame baseline on the held-out sequence.

Paper

Full text

PDF

Synergistic Fusion for Traversability: Combining Footsteps, Semantics, and Geometry

Semantic Scholar · 2025

Abstract

Acquiring dense, pixel-wise traversability labels for quadruped robots is a significant bottleneck that relies heavily on costly and time-consuming manual annotation, limiting the scalability and adaptability of robots. To address this, we propose a novel autolabeling pipeline that generates high-quality traversability labels without human intervention. Our method synergistically fuses three key sources of information available to the robot: (1) weak supervision from the robot’s own future footstep plans, derived from SLAM, to guide a video-native segmentation model; (2) closed-set semantic context from a Mask2Former model to disambiguate surfaces; and (3) geometric cues from an RGB-D sensor to ensure sharp boundary alignment. Through experiments conducted in diverse urban and natural environments in Daejeon, South Korea, we demonstrate that a lightweight segmentation model trained exclusively on our automatically generated labels significantly outperforms strong baselines, including state-of-the-art models pre-trained on large-scale public datasets like Mapillary Vistas. Notably, our approach mitigates domain shift and achieves a 0.11 higher mIoU than a SAM-prompted single-frame baseline on the held-out sequence.

Similar papers

© 2026 NYSGPT2525 LLC