Learning Object-Action Relations from Bimanual Human Demonstration Using Graph Networks

Recognizing human actions is a vital task for a humanoid robot, especially in\ndomains like programming by demonstration. Previous approaches on action\nrecognition primarily focused on the overall prevalent action being executed,\nbut we argue that bimanual human motion cannot always be described sufficiently\nwith a single action label. We present a system for frame-wise action\nclassification and segmentation in bimanual human demonstrations. The system\nextracts symbolic spatial object relations from raw RGB-D video data captured\nfrom the robot's point of view in order to build graph-based scene\nrepresentations. To learn object-action relations, a graph network classifier\nis trained using these representations together with ground truth action labels\nto predict the action executed by each hand.\n We evaluated the proposed classifier on a new RGB-D video dataset showing\ndaily action sequences focusing on bimanual manipulation actions. It consists\nof 6 subjects performing 9 tasks with 10 repetitions each, which leads to 540\nvideo recordings with 2 hours and 18 minutes total playtime and per-hand ground\ntruth action labels for each frame. We show that the classifier is able to\nreliably identify (action classification macro F1-score of 0.86) the true\nexecuted action of each hand within its top 3 predictions on a frame-by-frame\nbasis without prior temporal action segmentation.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC