Detecting Adversarial Examples and Other Misclassifications in Neural Networks by Introspection

Despite having excellent performances for a wide variety of tasks, modern\nneural networks are unable to provide a reliable confidence value allowing to\ndetect misclassifications. This limitation is at the heart of what is known as\nan adversarial example, where the network provides a wrong prediction\nassociated with a strong confidence to a slightly modified image. Moreover,\nthis overconfidence issue has also been observed for regular errors and\nout-of-distribution data. We tackle this problem by what we call introspection,\ni.e. using the information provided by the logits of an already pretrained\nneural network. We show that by training a simple 3-layers neural network on\ntop of the logit activations, we are able to detect misclassifications at a\ncompetitive level.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC