T-Miner: A Generative Approach to Defend Against Trojan Attacks on DNN-based Text Classification
Deep Neural Network (DNN) classifiers are known to be vulnerable to Trojan or\nbackdoor attacks, where the classifier is manipulated such that it\nmisclassifies any input containing an attacker-determined Trojan trigger.\nBackdoors compromise a model's integrity, thereby posing a severe threat to the\nlandscape of DNN-based classification. While multiple defenses against such\nattacks exist for classifiers in the image domain, there have been limited\nefforts to protect classifiers in the text domain.\n We present Trojan-Miner (T-Miner) -- a defense framework for Trojan attacks\non DNN-based text classifiers. T-Miner employs a sequence-to-sequence\n(seq-2-seq) generative model that probes the suspicious classifier and learns\nto produce text sequences that are likely to contain the Trojan trigger.\nT-Miner then analyzes the text produced by the generative model to determine if\nthey contain trigger phrases, and correspondingly, whether the tested\nclassifier has a backdoor. T-Miner requires no access to the training dataset\nor clean inputs of the suspicious classifier, and instead uses synthetically\ncrafted "nonsensical" text inputs to train the generative model. We extensively\nevaluate T-Miner on 1100 model instances spanning 3 ubiquitous DNN model\narchitectures, 5 different classification tasks, and a variety of trigger\nphrases. We show that T-Miner detects Trojan and clean models with a 98.75%\noverall accuracy, while achieving low false positives on clean models. We also\nshow that T-Miner is robust against a variety of targeted, advanced attacks\nfrom an adaptive attacker.\n
Paper
References (65)
Scroll for more · 38 remaining