There has been increasing interest in characterising the error behaviour of\nsystems which contain deep learning models before deploying them into any\nsafety-critical scenario. However, characterising such behaviour usually\nrequires large-scale testing of the model that can be extremely computationally\nexpensive for complex real-world tasks. For example, tasks involving compute\nintensive object detectors as one of their components. In this work, we propose\nan approach that enables efficient large-scale testing using simplified\nlow-fidelity simulators and without the computational cost of executing\nexpensive deep learning models. Our approach relies on designing an efficient\nsurrogate model corresponding to the compute intensive components of the task\nunder test. We demonstrate the efficacy of our methodology by evaluating the\nperformance of an autonomous driving task in the Carla simulator with reduced\ncomputational expense by training efficient surrogate models for PIXOR and\nCenterPoint LiDAR detectors, whilst demonstrating that the accuracy of the\nsimulation is maintained.\n