Multimodal representation models for prediction and control from partial information

Similar to humans, robots benefit from interacting with their environment\nthrough a number of different sensor modalities, such as vision, touch, sound.\nHowever, learning from different sensor modalities is difficult, because the\nlearning model must be able to handle diverse types of signals, and learn a\ncoherent representation even when parts of the sensor inputs are missing. In\nthis paper, a multimodal variational autoencoder is proposed to enable an iCub\nhumanoid robot to learn representations of its sensorimotor capabilities from\ndifferent sensor modalities. The proposed model is able to (1) reconstruct\nmissing sensory modalities, (2) predict the sensorimotor state of self and the\nvisual trajectories of other agents actions, and (3) control the agent to\nimitate an observed visual trajectory. Also, the proposed multimodal\nvariational autoencoder can capture the kinematic redundancy of the robot\nmotion through the learned probability distribution. Training multimodal models\nis not trivial due to the combinatorial complexity given by the possibility of\nmissing modalities. We propose a strategy to train multimodal models, which\nsuccessfully achieves improved performance of different reconstruction models.\nFinally, extensive experiments have been carried out using an iCub humanoid\nrobot, showing high performance in multiple reconstruction, prediction and\nimitation tasks.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC