Rel3D: A Minimally Contrastive Benchmark for Grounding Spatial Relations in 3D

Understanding spatial relations (e.g., "laptop on table") in visual input is\nimportant for both humans and robots. Existing datasets are insufficient as\nthey lack large-scale, high-quality 3D ground truth information, which is\ncritical for learning spatial relations. In this paper, we fill this gap by\nconstructing Rel3D: the first large-scale, human-annotated dataset for\ngrounding spatial relations in 3D. Rel3D enables quantifying the effectiveness\nof 3D information in predicting spatial relations on large-scale human data.\nMoreover, we propose minimally contrastive data collection -- a novel\ncrowdsourcing method for reducing dataset bias. The 3D scenes in our dataset\ncome in minimally contrastive pairs: two scenes in a pair are almost identical,\nbut a spatial relation holds in one and fails in the other. We empirically\nvalidate that minimally contrastive examples can diagnose issues with current\nrelation detection models as well as lead to sample-efficient training. Code\nand data are available at https://github.com/princeton-vl/Rel3D.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC