Exploring Temporal Dependencies in Multimodal Referring Expressions with Mixed Reality

In collaborative tasks, people rely both on verbal and non-verbal cues\nsimultaneously to communicate with each other. For human-robot interaction to\nrun smoothly and naturally, a robot should be equipped with the ability to\nrobustly disambiguate referring expressions. In this work, we propose a model\nthat can disambiguate multimodal fetching requests using modalities such as\nhead movements, hand gestures, and speech. We analysed the acquired data from\nmixed reality experiments and formulated a hypothesis that modelling temporal\ndependencies of events in these three modalities increases the model's\npredictive power. We evaluated our model on a Bayesian framework to interpret\nreferring expressions with and without exploiting a temporal prior.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC