There has been a recent surge in methods that aim to decompose and segment\nscenes into multiple objects in an unsupervised manner, i.e., unsupervised\nmulti-object segmentation. Performing such a task is a long-standing goal of\ncomputer vision, offering to unlock object-level reasoning without requiring\ndense annotations to train segmentation models. Despite significant progress,\ncurrent models are developed and trained on visually simple scenes depicting\nmono-colored objects on plain backgrounds. The natural world, however, is\nvisually complex with confounding aspects such as diverse textures and\ncomplicated lighting effects. In this study, we present a new benchmark called\nClevrTex, designed as the next challenge to compare, evaluate and analyze\nalgorithms. ClevrTex features synthetic scenes with diverse shapes, textures\nand photo-mapped materials, created using physically based rendering\ntechniques. It includes 50k examples depicting 3-10 objects arranged on a\nbackground, created using a catalog of 60 materials, and a further test set\nfeaturing 10k images created using 25 different materials. We benchmark a large\nset of recent unsupervised multi-object segmentation models on ClevrTex and\nfind all state-of-the-art approaches fail to learn good representations in the\ntextured setting, despite impressive performance on simpler data. We also\ncreate variants of the ClevrTex dataset, controlling for different aspects of\nscene complexity, and probe current approaches for individual shortcomings.\nDataset and code are available at\nhttps://www.robots.ox.ac.uk/~vgg/research/clevrtex.\n