A Variational Graph Autoencoder for Manipulation Action Recognition and Prediction

Despite decades of research, understanding human manipulation activities is,\nand has always been, one of the most attractive and challenging research topics\nin computer vision and robotics. Recognition and prediction of observed human\nmanipulation actions have their roots in the applications related to, for\ninstance, human-robot interaction and robot learning from demonstration. The\ncurrent research trend heavily relies on advanced convolutional neural networks\nto process the structured Euclidean data, such as RGB camera images. These\nnetworks, however, come with immense computational complexity to be able to\nprocess high dimensional raw data.\n Different from the related works, we here introduce a deep graph autoencoder\nto jointly learn recognition and prediction of manipulation tasks from symbolic\nscene graphs, instead of relying on the structured Euclidean data. Our network\nhas a variational autoencoder structure with two branches: one for identifying\nthe input graph type and one for predicting the future graphs. The input of the\nproposed network is a set of semantic graphs which store the spatial relations\nbetween subjects and objects in the scene. The network output is a label set\nrepresenting the detected and predicted class types. We benchmark our new model\nagainst different state-of-the-art methods on two different datasets, MANIAC\nand MSRC-9, and show that our proposed model can achieve better performance. We\nalso release our source code https://github.com/gamzeakyol/GNet.\n

Paper

References (41)

Scroll for more · 29 remaining

Similar papers

© 2026 NYSGPT2525 LLC