VoRTX: Volumetric 3D Reconstruction With Transformers for Voxelwise View Selection and Fusion

Recent volumetric 3D reconstruction methods can produce very accurate\nresults, with plausible geometry even for unobserved surfaces. However, they\nface an undesirable trade-off when it comes to multi-view fusion. They can fuse\nall available view information by global averaging, thus losing fine detail, or\nthey can heuristically cluster views for local fusion, thus restricting their\nability to consider all views jointly. Our key insight is that greater detail\ncan be retained without restricting view diversity by learning a view-fusion\nfunction conditioned on camera pose and image content. We propose to learn this\nmulti-view fusion using a transformer. To this end, we introduce VoRTX, an\nend-to-end volumetric 3D reconstruction network using transformers for\nwide-baseline, multi-view feature fusion. Our model is occlusion-aware,\nleveraging the transformer architecture to predict an initial, projective scene\ngeometry estimate. This estimate is used to avoid backprojecting image features\nthrough surfaces into occluded regions. We train our model on ScanNet and show\nthat it produces better reconstructions than state-of-the-art methods. We also\ndemonstrate generalization without any fine-tuning, outperforming the same\nstate-of-the-art methods on two other datasets, TUM-RGBD and ICL-NUIM.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC