With the increasing adoption of Deep Learning (DL) for critical tasks, such\nas autonomous driving, the evaluation of the quality of systems that rely on DL\nhas become crucial. Once trained, DL systems produce an output for any\narbitrary numeric vector provided as input, regardless of whether it is within\nor outside the validity domain of the system under test. Hence, the quality of\nsuch systems is determined by the intersection between their validity domain\nand the regions where their outputs exhibit a misbehaviour. In this paper, we\nintroduce the notion of frontier of behaviours, i.e., the inputs at which the\nDL system starts to misbehave. If the frontier of misbehaviours is outside the\nvalidity domain of the system, the quality check is passed. Otherwise, the\ninputs at the intersection represent quality deficiencies of the system. We\ndeveloped DeepJanus, a search-based tool that generates frontier inputs for DL\nsystems. The experimental results obtained for the lane keeping component of a\nself-driving car show that the frontier of a well trained system contains\nalmost exclusively unrealistic roads that violate the best practices of civil\nengineering, while the frontier of a poorly trained one includes many valid\ninputs that point to serious deficiencies of the system.\n
Paper
References (56)
Scroll for more · 38 remaining