OmniDet: Surround View Cameras based Multi-task Visual Perception Network for Autonomous Driving

Surround View fisheye cameras are commonly deployed in automated driving for\n360\\deg{} near-field sensing around the vehicle. This work presents a\nmulti-task visual perception network on unrectified fisheye images to enable\nthe vehicle to sense its surrounding environment. It consists of six primary\ntasks necessary for an autonomous driving system: depth estimation, visual\nodometry, semantic segmentation, motion segmentation, object detection, and\nlens soiling detection. We demonstrate that the jointly trained model performs\nbetter than the respective single task versions. Our multi-task model has a\nshared encoder providing a significant computational advantage and has\nsynergized decoders where tasks support each other. We propose a novel camera\ngeometry based adaptation mechanism to encode the fisheye distortion model both\nat training and inference. This was crucial to enable training on the WoodScape\ndataset, comprised of data from different parts of the world collected by 12\ndifferent cameras mounted on three different cars with different intrinsics and\nviewpoints. Given that bounding boxes is not a good representation for\ndistorted fisheye images, we also extend object detection to use a polygon with\nnon-uniformly sampled vertices. We additionally evaluate our model on standard\nautomotive datasets, namely KITTI and Cityscapes. We obtain the\nstate-of-the-art results on KITTI for depth estimation and pose estimation\ntasks and competitive performance on the other tasks. We perform extensive\nablation studies on various architecture choices and task weighting\nmethodologies. A short video at https://youtu.be/xbSjZ5OfPes provides\nqualitative results.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC