Automatically Determining Poisonous Attacks on Neural Networks

Patent №

US 11,645,515

Granted

2023-05-09

Filed 2019

Owner

INTERNATIONAL BUSINESS MACHINES CORPORATION

AI components

5

ml · nlp · vision · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16571323

Embodiments relate to a system, program product, and method for automatically determining which activation data points in a neural model have been poisoned to erroneously indicate association with a particular label or labels. A neural network is trained using potentially poisoned training data. Each of the training data points is classified using the network to retain the activations of the last hidden layer, and segment those activations by the label of corresponding training data. Clustering is applied to the retained activations of each segment, and a cluster assessment is conducted for each cluster associated with each label to distinguish clusters with potentially poisoned activations from clusters populated with legitimate activations. The assessment includes executing a set of analyses and integrating the results of the analyses into a determination as to whether a training data set is poisonous based on determining if resultant activation clusters are poisoned.

Machine learningNatural languageVisionKnowledge representationAI hardwareG06N 3/08G06F 18/211G06F 18/23G06F 18/24G06N 3/0499G06N 3/09G06N 7/01G06N 20/00+3 more

AI classification

Machine learning1.00
Vision1.00
Natural language1.00
AI hardware1.00
Knowledge representation1.00
Planning0.04
Speech0.01
Evolutionary computation0.00

Ownership

INTERNATIONAL BUSINESS MACHINES CORPORATION

assignment · 503800501

Assignors

ANGEL, NATHALIE BARACALDO, CHEN, BRYANT, SRIVASTAVA, BIPLAV, LUDWIG, HEIKO H.

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC