Deep generative models allow for photorealistic image synthesis at high\nresolutions. But for many applications, this is not enough: content creation\nalso needs to be controllable. While several recent works investigate how to\ndisentangle underlying factors of variation in the data, most of them operate\nin 2D and hence ignore that our world is three-dimensional. Further, only few\nworks consider the compositional nature of scenes. Our key hypothesis is that\nincorporating a compositional 3D scene representation into the generative model\nleads to more controllable image synthesis. Representing scenes as\ncompositional generative neural feature fields allows us to disentangle one or\nmultiple objects from the background as well as individual objects' shapes and\nappearances while learning from unstructured and unposed image collections\nwithout any additional supervision. Combining this scene representation with a\nneural rendering pipeline yields a fast and realistic image synthesis model. As\nevidenced by our experiments, our model is able to disentangle individual\nobjects and allows for translating and rotating them in the scene as well as\nchanging the camera pose.\n