MaskFusion: Real-Time Recognition, Tracking and Reconstruction of Multiple Moving Objects

We present MaskFusion, a real-time, object-aware, semantic and dynamic RGB-D\nSLAM system that goes beyond traditional systems which output a purely\ngeometric map of a static scene. MaskFusion recognizes, segments and assigns\nsemantic class labels to different objects in the scene, while tracking and\nreconstructing them even when they move independently from the camera.\n As an RGB-D camera scans a cluttered scene, image-based instance-level\nsemantic segmentation creates semantic object masks that enable real-time\nobject recognition and the creation of an object-level representation for the\nworld map. Unlike previous recognition-based SLAM systems, MaskFusion does not\nrequire known models of the objects it can recognize, and can deal with\nmultiple independent motions. MaskFusion takes full advantage of using\ninstance-level semantic segmentation to enable semantic labels to be fused into\nan object-aware map, unlike recent semantics enabled SLAM systems that perform\nvoxel-level semantic segmentation. We show augmented-reality applications that\ndemonstrate the unique features of the map output by MaskFusion:\ninstance-aware, semantic and dynamic.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC