Semantic Implicit Neural Scene Representations With Semi-Supervised Training

The recent success of implicit neural scene representations has presented a\nviable new method for how we capture and store 3D scenes. Unlike conventional\n3D representations, such as point clouds, which explicitly store scene\nproperties in discrete, localized units, these implicit representations encode\na scene in the weights of a neural network which can be queried at any\ncoordinate to produce these same scene properties. Thus far, implicit\nrepresentations have primarily been optimized to estimate only the appearance\nand/or 3D geometry information in a scene. We take the next step and\ndemonstrate that an existing implicit representation (SRNs) is actually\nmulti-modal; it can be further leveraged to perform per-point semantic\nsegmentation while retaining its ability to represent appearance and geometry.\nTo achieve this multi-modal behavior, we utilize a semi-supervised learning\nstrategy atop the existing pre-trained scene representation. Our method is\nsimple, general, and only requires a few tens of labeled 2D segmentation masks\nin order to achieve dense 3D semantic segmentation. We explore two novel\napplications for this semantically aware implicit neural scene representation:\n3D novel view and semantic label synthesis given only a single input RGB image\nor 2D label mask, as well as 3D interpolation of appearance and semantics.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC