UniNet: A Unified Scene Understanding Network and Exploring Multi-Task Relationships through the Lens of Adversarial Attacks
Scene understanding is crucial for autonomous systems which intend to operate\nin the real world. Single task vision networks extract information only based\non some aspects of the scene. In multi-task learning (MTL), on the other hand,\nthese single tasks are jointly learned, thereby providing an opportunity for\ntasks to share information and obtain a more comprehensive understanding. To\nthis end, we develop UniNet, a unified scene understanding network that\naccurately and efficiently infers vital vision tasks including object\ndetection, semantic segmentation, instance segmentation, monocular depth\nestimation, and monocular instance depth prediction. As these tasks look at\ndifferent semantic and geometric information, they can either complement or\nconflict with each other. Therefore, understanding inter-task relationships can\nprovide useful cues to enable complementary information sharing. We evaluate\nthe task relationships in UniNet through the lens of adversarial attacks based\non the notion that they can exploit learned biases and task interactions in the\nneural network. Extensive experiments on the Cityscapes dataset, using\nuntargeted and targeted attacks reveal that semantic tasks strongly interact\namongst themselves, and the same holds for geometric tasks. Additionally, we\nshow that the relationship between semantic and geometric tasks is asymmetric\nand their interaction becomes weaker as we move towards higher-level\nrepresentations.\n