The MECCANO Dataset: Understanding Human-Object Interactions from Egocentric Videos in an Industrial-like Domain

Wearable cameras allow to collect images and videos of humans interacting\nwith the world. While human-object interactions have been thoroughly\ninvestigated in third person vision, the problem has been understudied in\negocentric settings and in industrial scenarios. To fill this gap, we introduce\nMECCANO, the first dataset of egocentric videos to study human-object\ninteractions in industrial-like settings. MECCANO has been acquired by 20\nparticipants who were asked to build a motorbike model, for which they had to\ninteract with tiny objects and tools. The dataset has been explicitly labeled\nfor the task of recognizing human-object interactions from an egocentric\nperspective. Specifically, each interaction has been labeled both temporally\n(with action segments) and spatially (with active object bounding boxes). With\nthe proposed dataset, we investigate four different tasks including 1) action\nrecognition, 2) active object detection, 3) active object recognition and 4)\negocentric human-object interaction detection, which is a revisited version of\nthe standard human-object interaction detection task. Baseline results show\nthat the MECCANO dataset is a challenging benchmark to study egocentric\nhuman-object interactions in industrial-like scenarios. We publicy release the\ndataset at https://iplab.dmi.unict.it/MECCANO.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC