Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D Scans

This paper introduces a pipeline to parametrically sample and render\nmulti-task vision datasets from comprehensive 3D scans from the real world.\nChanging the sampling parameters allows one to "steer" the generated datasets\nto emphasize specific information. In addition to enabling interesting lines of\nresearch, we show the tooling and generated data suffice to train robust vision\nmodels.\n Common architectures trained on a generated starter dataset reached\nstate-of-the-art performance on multiple common vision tasks and benchmarks,\ndespite having seen no benchmark or non-pipeline data. The depth estimation\nnetwork outperforms MiDaS and the surface normal estimation network is the\nfirst to achieve human-level performance for in-the-wild surface normal\nestimation -- at least according to one metric on the OASIS benchmark.\n The Dockerized pipeline with CLI, the (mostly python) code, PyTorch\ndataloaders for the generated data, the generated starter dataset, download\nscripts and other utilities are available through our project website,\nhttps://omnidata.vision.\n

Paper

References (74)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC