Advances in unsupervised learning of object-representations have culminated\nin the development of a broad range of methods for unsupervised object\nsegmentation and interpretable object-centric scene generation. These methods,\nhowever, are limited to simulated and real-world datasets with limited visual\ncomplexity. Moreover, object representations are often inferred using RNNs\nwhich do not scale well to large images or iterative refinement which avoids\nimposing an unnatural ordering on objects in an image but requires the a priori\ninitialisation of a fixed number of object representations. In contrast to\nestablished paradigms, this work proposes an embedding-based approach in which\nembeddings of pixels are clustered in a differentiable fashion using a\nstochastic stick-breaking process. Similar to iterative refinement, this\nclustering procedure also leads to randomly ordered object representations, but\nwithout the need of initialising a fixed number of clusters a priori. This is\nused to develop a new model, GENESIS-v2, which can infer a variable number of\nobject representations without using RNNs or iterative refinement. We show that\nGENESIS-v2 performs strongly in comparison to recent baselines in terms of\nunsupervised image segmentation and object-centric scene generation on\nestablished synthetic datasets as well as more complex real-world datasets.\n