Chained Representation Cycling: Learning to Estimate 3D Human Pose and Shape by Cycling Between Representations

The goal of many computer vision systems is to transform image pixels into 3D\nrepresentations. Recent popular models use neural networks to regress directly\nfrom pixels to 3D object parameters. Such an approach works well when\nsupervision is available, but in problems like human pose and shape estimation,\nit is difficult to obtain natural images with 3D ground truth. To go one step\nfurther, we propose a new architecture that facilitates unsupervised, or\nlightly supervised, learning. The idea is to break the problem into a series of\ntransformations between increasingly abstract representations. Each step\ninvolves a cycle designed to be learnable without annotated training data, and\nthe chain of cycles delivers the final solution. Specifically, we use 2D body\npart segments as an intermediate representation that contains enough\ninformation to be lifted to 3D, and at the same time is simple enough to be\nlearned in an unsupervised way. We demonstrate the method by learning 3D human\npose and shape from un-paired and un-annotated images. We also explore varying\namounts of paired data and show that cycling greatly alleviates the need for\npaired data. While we present results for modeling humans, our formulation is\ngeneral and can be applied to other vision problems.\n

Paper

References (43)

Scroll for more · 31 remaining

Similar papers

© 2026 NYSGPT2525 LLC