The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and Pose

The availability of a large labeled dataset is a key requirement for applying\ndeep learning methods to solve various computer vision tasks. In the context of\nunderstanding human activities, existing public datasets, while large in size,\nare often limited to a single RGB camera and provide only per-frame or per-clip\naction annotations. To enable richer analysis and understanding of human\nactivities, we introduce IKEA ASM -- a three million frame, multi-view,\nfurniture assembly video dataset that includes depth, atomic actions, object\nsegmentation, and human pose. Additionally, we benchmark prominent methods for\nvideo action recognition, object segmentation and human pose estimation tasks\non this challenging dataset. The dataset enables the development of holistic\nmethods, which integrate multi-modal and multi-view data to better perform on\nthese tasks.\n

Paper

References (77)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC