VPFusion: Joint 3D Volume and Pixel-Aligned Feature Fusion for Single and Multi-view 3D Reconstruction
We introduce a unified single and multi-view neural implicit 3D\nreconstruction framework VPFusion. VPFusion attains high-quality reconstruction\nusing both - 3D feature volume to capture 3D-structure-aware context, and\npixel-aligned image features to capture fine local detail. Existing approaches\nuse RNN, feature pooling, or attention computed independently in each view for\nmulti-view fusion. RNNs suffer from long-term memory loss and permutation\nvariance, while feature pooling or independently computed attention leads to\nrepresentation in each view being unaware of other views before the final\npooling step. In contrast, we show improved multi-view feature fusion by\nestablishing transformer-based pairwise view association. In particular, we\npropose a novel interleaved 3D reasoning and pairwise view association\narchitecture for feature volume fusion across different views. Using this\nstructure-aware and multi-view-aware feature volume, we show improved 3D\nreconstruction performance compared to existing methods. VPFusion improves the\nreconstruction quality further by also incorporating pixel-aligned local image\nfeatures to capture fine detail. We verify the effectiveness of VPFusion on the\nShapeNet and ModelNet datasets, where we outperform or perform on-par the\nstate-of-the-art single and multi-view 3D shape reconstruction methods.\n