For many fundamental scene understanding tasks, it is difficult or impossible\nto obtain per-pixel ground truth labels from real images. We address this\nchallenge by introducing Hypersim, a photorealistic synthetic dataset for\nholistic indoor scene understanding. To create our dataset, we leverage a large\nrepository of synthetic scenes created by professional artists, and we generate\n77,400 images of 461 indoor scenes with detailed per-pixel labels and\ncorresponding ground truth geometry. Our dataset: (1) relies exclusively on\npublicly available 3D assets; (2) includes complete scene geometry, material\ninformation, and lighting information for every scene; (3) includes dense\nper-pixel semantic instance segmentations and complete camera information for\nevery image; and (4) factors every image into diffuse reflectance, diffuse\nillumination, and a non-diffuse residual term that captures view-dependent\nlighting effects.\n We analyze our dataset at the level of scenes, objects, and pixels, and we\nanalyze costs in terms of money, computation time, and annotation effort.\nRemarkably, we find that it is possible to generate our entire dataset from\nscratch, for roughly half the cost of training a popular open-source natural\nlanguage processing model. We also evaluate sim-to-real transfer performance on\ntwo real-world scene understanding tasks - semantic segmentation and 3D shape\nprediction - where we find that pre-training on our dataset significantly\nimproves performance on both tasks, and achieves state-of-the-art performance\non the most challenging Pix3D test set. All of our rendered image data, as well\nas all the code we used to generate our dataset and perform our experiments, is\navailable online.\n