Learning Object Placements For Relational Instructions by Hallucinating Scene Representations

Robots coexisting with humans in their environment and performing services\nfor them need the ability to interact with them. One particular requirement for\nsuch robots is that they are able to understand spatial relations and can place\nobjects in accordance with the spatial relations expressed by their user. In\nthis work, we present a convolutional neural network for estimating pixelwise\nobject placement probabilities for a set of spatial relations from a single\ninput image. During training, our network receives the learning signal by\nclassifying hallucinated high-level scene representations as an auxiliary task.\nUnlike previous approaches, our method does not require ground truth data for\nthe pixelwise relational probabilities or 3D models of the objects, which\nsignificantly expands the applicability in practical applications. Our results\nobtained using real-world data and human-robot experiments demonstrate the\neffectiveness of our method in reasoning about the best way to place objects to\nreproduce a spatial relation. Videos of our experiments can be found at\nhttps://youtu.be/zaZkHTWFMKM\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC