Towards Optimal Strategies for Training Self-Driving Perception Models in Simulation

Autonomous driving relies on a huge volume of real-world data to be labeled\nto high precision. Alternative solutions seek to exploit driving simulators\nthat can generate large amounts of labeled data with a plethora of content\nvariations. However, the domain gap between the synthetic and real data\nremains, raising the following important question: What are the best ways to\nutilize a self-driving simulator for perception tasks? In this work, we build\non top of recent advances in domain-adaptation theory, and from this\nperspective, propose ways to minimize the reality gap. We primarily focus on\nthe use of labels in the synthetic domain alone. Our approach introduces both a\nprincipled way to learn neural-invariant representations and a theoretically\ninspired view on how to sample the data from the simulator. Our method is easy\nto implement in practice as it is agnostic of the network architecture and the\nchoice of the simulator. We showcase our approach on the bird's-eye-view\nvehicle segmentation task with multi-sensor data (cameras, lidar) using an\nopen-source simulator (CARLA), and evaluate the entire framework on a\nreal-world dataset (nuScenes). Last but not least, we show what types of\nvariations (e.g. weather conditions, number of assets, map design, and color\ndiversity) matter to perception networks when trained with driving simulators,\nand which ones can be compensated for with our domain adaptation technique.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC