ARTIC3D: Learning Robust Articulated 3D Shapes from Noisy Web Image Collections

Estimating 3D articulated shapes like animal bodies from monocular images is inherently challenging due to the ambiguities of camera viewpoint, pose, texture, lighting, etc. We propose ARTIC3D, a self-supervised framework to reconstruct per-instance 3D shapes from a sparse image collection in-the-wild. Specifically, ARTIC3D is built upon a skeleton-based surface representation and is further guided by 2D diffusion priors from Stable Diffusion. First, we enhance the input images with occlusions/truncation via 2D diffusion to obtain cleaner mask estimates and semantic features. Second, we perform diffusion-guided 3D optimization to estimate shape and texture that are of high-fidelity and faithful to input images. We also propose a novel technique to calculate more stable image-level gradients via diffusion models compared to existing alternatives. Finally, we produce realistic animations by fine-tuning the rendered shape and texture under rigid part transformations. Extensive evaluations on multiple existing datasets as well as newly introduced noisy web image collections with occlusions and truncation demonstrate that ARTIC3D outputs are more robust to noisy images, higher quality in terms of shape and texture details, and more realistic when animated. Project page: https://chhankyao.github.io/artic3d/

Paper

References (46)

Scroll for more · 34 remaining

Similar papers

Peer review

Reviewer atAr5/10 · confidence 4/52023-07-04

Summary

ARTIC3D is a self-supervised framework that reconstructs 3D articulated shapes and textures of animals from sparse and noisy online images. It uses a skeleton-based surface representation and 2D diffusion priors to enhance the input images and guide the 3D optimization. It also enables realistic animation by fine-tuning the rendered shape and texture under rigid part transformations. ARTIC3D outperforms prior methods in terms of shape and texture fidelity, robustness to occlusions and truncation, and pose transferability. The authors also introduce E-LASSIE, an extended dataset with noisy web images, to evaluate model robustness.

Strengths

* ARTIC3D is a self-supervised framework that can reconstruct 3D articulated shapes and textures of animals from sparse and noisy online images, without relying on any pre-defined shape templates or per-image annotations. This makes it scalable and adaptable to different animal species and poses. * The method leverages 2D diffusion to enhance the input images by removing occlusions and truncation, and to extract semantic features and 3D skeleton initialization. This improves the quality and robustness of the 3D outputs, as well as the efficiency and stability of the optimization process. * Usage of diffusion-guided 3D optimization to estimate shape and texture that are faithful to the input images and consistent across different viewpoints and poses. It also introduces a novel technique to calculate more stable image-level gradients via diffusion models, which enhances the convergence and robustness of the optimization. * ARTIC3D produces realistic animations by fine-tuning the rendered shape and texture under rigid part transformations, which preserves the articulation and details of the 3D shapes. It also enables explicit pose transfer and animation, which are not feasible for prior diffusion-guided methods with neural volumetric representations.

Weaknesses

* The manuscript is not easy to follow for anybody who is unfamiliar with LASSIE and HI-LASSIE. This work is a step forward from Hi-LASSIE using Stable Diffusion to avoid several pitfalls with respect to optimization and image pre-processing. * ARTIC3D depends on the 3D skeleton initialization from Hi-LASSIE [38], which may be inaccurate for occluded or truncated animal bodies, resulting in unrealistic part shapes. Also, it struggles with fluffy animals with ambiguous skeletal configuration, such as sheep, which pose challenges in skeleton discovery and shape reconstruction. * The reconstruction results are not significantly better than Hi-LASSIE in most case 1% better, however it is not clear from the manuscript whether this is statistically significant or not.

Questions

* The usage of SD improves texture quality on a perceptual level. However, as seen in many figures (ex 3, 4 in supp, 3 main) it mainly hallucinates the texture to appear realistic however further apart from the image used as reference. For example, the reconstructed elephant looks nothing like the image, similarly the kangaroo, tiger, zebra. The work is presented as reconstruction work and as such diverging from the images can't be considered a proper reconstruction. Have you examined a way to mitigate this texture drift? * The user studies on animation are flawed. A 55% preference from 100 users means that ARTIC3D'a animations are only slightly better than rigid transform which is close enough to random choice. Could you explain in detail what sample size was used for the user study? Whether the studies were cherry picked prior to the user study? I find the user study explanation to lack a lot of key details.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

2 fair

Limitations

Yes

Authorsrebuttal2023-08-14

Please let us know whether you have additional questions after reading our response

We appreciate your reviews and comments. We hope our responses address your concerns. Please let us know if you have further questions after reading our rebuttal. We hope to address all the potential issues during the discussion period.

Authorsrebuttal2023-08-17

Please let us know whether all questions have been addressed

Dear Reviewer, As we are approaching the midpoint of the discussion period, we would like to confirm whether we have successfully addressed the raised concerns in your review. Should any lingering issues require further attention, please let us know as early as possible so we can answer them soon. We appreciate your time and effort in enhancing the quality of our manuscript. Thank you

Reviewer GaMw5/10 · confidence 4/52023-07-05

Summary

This paper proposes a method to reconstruct the shape and texture of articulated objects from noisy web image collections. To achieve this, ARTIC3D proposed a diffusion-based 2D image enhancement module DASS, and then reconstructed the shape and texture maps using Hi-LASSIE. Moreover, to increase the animation results, ARTIC3D introduced a T-DASS module for animation fine-tuning. Experiments on the E-LASSIE dataset show that this method can produce high-fidelity animation results from noisy inputs.

Strengths

--ARTIC3D can directly reconstruct the shapes and texture maps of articulated objects from noisy web image collections, greetly increase the robotness of Hi-LASSIE. --The proposed DASS and T-DASS modules are novel, intuitive and effective. --The paper is well-writen and easy to follow.

Weaknesses

--The texture map obtained by ARTIC3D is not good enough. To solve this problem, ARTIC3D relies on the diffusion-based DASS module and the animation fine-tuning module to achieve high-fidelity animation and novel view rendering results. However, this may lead to 3D inconsistency. How does this method solve this problem? --Although enhanced by T-DASS, the animation results are still blurry and the texture moves over time. Also, the T-DASS module seems to make the results blurrier than directly using DASS module. --The reconstructed texture and shape are not faithful to the input image in some cases. The DASS module change the shape and appearance of the input images. Moreover, the images generated by this module may lose their 3D consistency.

Questions

--ARTIC3D optimized the textured images per instance(L220-211). However, the appearance of a single animal may also vary with the lighting conditions. How does this method deal with this problem? --ARTIC3D ustilizes the T-DASS module to enhance animation. It would be great if the paper can also include some discussions on rendering speed and other relevant computational costs. --The T-DASS module may lead to 3D inconsistency. How does this method solve this problem?

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

--The shapes and textures generated by this method is inaccurate and blurry, and are not faithful to the input images in some cases. --The DASS module may lead to 3D inconsistency.

Authorsrebuttal2023-08-14

Please let us know whether you have additional questions after reading our response

We appreciate your reviews and comments. We hope our responses address your concerns. Please let us know if you have further questions after reading our rebuttal. We hope to address all the potential issues during the discussion period.

Authorsrebuttal2023-08-17

Please let us know whether all questions have been addressed

Dear Reviewer, As we are approaching the midpoint of the discussion period, we would like to confirm whether we have successfully addressed the raised concerns in your review. Should any lingering issues require further attention, please let us know as early as possible so we can answer them soon. We appreciate your time and effort in enhancing the quality of our manuscript. Thank you

Reviewer V4Hg6/10 · confidence 5/52023-07-06

Summary

This paper proposes a new framework, named ARTIC3D, to address the task of 3D reconstruction of articulated shapes and texture from noisy and few images. It is based on pre-trained diffusion models. Specifically, the authors use a novel decoder-based accumulative score sampling (DASS) to replace score distillation sampling (SDS). This has been used in many other frameworks to calculate pixel gradients in order to use diffusion priors more efficiently. Besides, they also extend the LASSIE dataset with more annotated data, called E-LASSIE, which could be useful for future works , especially for robustness of 3D reconstruction models.

Strengths

(1) Good writing. (2) Plentiful experiments to prove the effectiveness of their framework and necessity of each module. (3) Novelty of the method (DASS) to better implement 2D diffusion priors in 3D reconstruction tasks, especially when big and clean datasets are not available. Detailed techniques are elaborated in the method section, including how to use them in preprocessing noisy input image, shape and texture optimization, and animation fine-tuning. (4) Extension of the LASSIE dataset with more annotated images to a new dataset, E-LASSIE, which could be useful for future research in this area, especially for evaluating the 3D reconstruction model robustness.

Weaknesses

(1) This framework is highly based on LASSIE, both the method and the necessary input. Specifically, this framework needs the 3D skeleton from a pre-trained LASSIE model. And the framework also shares many parts with LASSIE's. But as mentioned in the Strengths part, the authors propose a new method to better incorporate diffusion models in 3D reconstruction and they propose a new dataset. (2) Though the title claims "learning from noisy web image collections", the noise actually only involves truncation and occlusions. There are also many other types of noise that have not been explored, including illumination variations, too small instances, multiple instances, etc., which are common in web images. They also uses DINO-VIT to do the foreground-background segmentation, which is acceptable but tricky, since the noisy background is also an important and common difficulty when dealing with web images.

Questions

n/a

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Limitations have already been elaborated clearly in their paper.

Authorsrebuttal2023-08-14

Please let us know whether you have additional questions after reading our response

We appreciate your reviews and comments. We hope our responses address your concerns. Please let us know if you have further questions after reading our rebuttal. We hope to address all the potential issues during the discussion period.

Authorsrebuttal2023-08-17

Please let us know whether all questions have been addressed

Dear Reviewer, As we are approaching the midpoint of the discussion period, we would like to confirm whether we have successfully addressed the raised concerns in your review. Should any lingering issues require further attention, please let us know as early as possible so we can answer them soon. We appreciate your time and effort in enhancing the quality of our manuscript. Thank you

Reviewer V4Hg2023-08-18

I thank the authors for the clarifications. I will keep my initial rating.

Reviewer MsPL6/10 · confidence 4/52023-07-07

Summary

This paper introduces an articulated 3D shape reconstruction method from noisy web images with the help of diffusion models. The authors use a diffusion method to enhance the noisy input images to get clean reference 2D images and masks. Then, skeleton-based surface representations are optimized from the reference images. Then, fine-tuning improves the animations of the reconstructed shapes.

Strengths

1. The overall narrative of the paper is sound and readable. 2. The authors propose to produce clean input reference images from noisy web images using diffusion models as a preprocessing step. 3. The authors propose Decoder-based Accumulative Score Sampling (DASS) to improve efficiency and reduce artifacts. 4. The authors designed a fine-tuning step to allow better animation of the reconstructed objects.

Weaknesses

1. The reconstruction part heavily relies on previous works like LASSIE[39] and Hi-LASSIE[38], it seems that the authors did not contribute very much to the core algorithm in this reconstruction task. Diffusion models helped improve the results, but most of the contribution is in the data cleaning and preprocessing. 2. Diffusion models help in image preprocessing, but Stable Diffusion would also produce results from its knowledge based on the text prompt. Therefore we can observe that some textures of the reconstructed objects are different from the reference input images, even for the unoccluded parts. 3. Figure 2 is not clear enough to show the whole workflow of the proposed methods, it is hard to relate the DASS module with the shape and texture optimization. 4. Lacking ablation studies, it is unclear whether the Distilling 3D reconstruction is working in section 3.4, without the Distilling 3D reconstruction part, the reconstruction technique would be mostly relied on LASSIE[39] and Hi-LASSIE[38] as I mentioned above.

Questions

1. How does the diffusion-guided optimization of shape and texture perform? It would be nice to see the ablation study on this part to distinguish this method from LASSIE[39] and Hi-LASSIE[38]. 2. How do you deal with the extra noise or overcorrection from the diffusion model? The preprocessed image may look like an average animal in the species from the diffusion models.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

The authors mentioned some limitations but not enough for me. Like the possible noise introduced from Stable Diffusion as mentioned above.

Authorsrebuttal2023-08-14

Please let us know whether you have additional questions after reading our response

We appreciate your reviews and comments. We hope our responses address your concerns. Please let us know if you have further questions after reading our rebuttal. We hope to address all the potential issues during the discussion period.

Reviewer MsPL2023-08-14

Thank the author for the detailed response. Most of my concerns are resolved. I would raise my rating to 6.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC