Learning monocular 3D reconstruction of articulated categories from motion

Monocular 3D reconstruction of articulated object categories is challenging\ndue to the lack of training data and the inherent ill-posedness of the problem.\nIn this work we use video self-supervision, forcing the consistency of\nconsecutive 3D reconstructions by a motion-based cycle loss. This largely\nimproves both optimization-based and learning-based 3D mesh reconstruction. We\nfurther introduce an interpretable model of 3D template deformations that\ncontrols a 3D surface through the displacement of a small number of local,\nlearnable handles. We formulate this operation as a structured layer relying on\nmesh-laplacian regularization and show that it can be trained in an end-to-end\nmanner. We finally introduce a per-sample numerical optimisation approach that\njointly optimises over mesh displacements and cameras within a video, boosting\naccuracy both for training and also as test time post-processing. While relying\nexclusively on a small set of videos collected per category for supervision, we\nobtain state-of-the-art reconstructions with diverse shapes, viewpoints and\ntextures for multiple articulated object categories.\n

Paper

References (62)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC