Inspired by the effectiveness of 3D Gaussian Splatting (3DGS) in reconstructing detailed 3D scenes within multiview setups and the emergence of large 2D human foundation models, we introduce Arc2Avatar, the first Score Distillation Sampling (SDS) based method utilizing a human face foundation model as guidance with just a single image as input. To achieve that, we extend such a model for diverse-view human head generation by fine-tuning on synthetic data and modifying its conditioning. Our avatars maintain a dense correspondence with a human face mesh template, allowing blendshape-based expression generation. This is achieved through a modified 3DGS approach, connectivity regularizers, and a strategic initialization tailored for our task. Additionally, we propose an optional efficient SDSbased correction step to refine the blendshape expressions. Experiments demonstrate that Arc2Avatar achieves stateof-the-art realism and identity preservation, effectively addressing color issues by allowing the use of very low guidance, enabled by our strong identity prior and initialization strategy, without compromising detail. Code and models are available on our project page.
Paper
References (100)
Scroll for more · 38 remaining