Adversarial Attacks on Black Box Video Classifiers: Leveraging the Power of Geometric Transformations

When compared to the image classification models, black-box adversarial\nattacks against video classification models have been largely understudied.\nThis could be possible because, with video, the temporal dimension poses\nsignificant additional challenges in gradient estimation. Query-efficient\nblack-box attacks rely on effectively estimated gradients towards maximizing\nthe probability of misclassifying the target video. In this work, we\ndemonstrate that such effective gradients can be searched for by parameterizing\nthe temporal structure of the search space with geometric transformations.\nSpecifically, we design a novel iterative algorithm Geometric TRAnsformed\nPerturbations (GEO-TRAP), for attacking video classification models. GEO-TRAP\nemploys standard geometric transformation operations to reduce the search space\nfor effective gradients into searching for a small group of parameters that\ndefine these operations. This group of parameters describes the geometric\nprogression of gradients, resulting in a reduced and structured search space.\nOur algorithm inherently leads to successful perturbations with surprisingly\nfew queries. For example, adversarial examples generated from GEO-TRAP have\nbetter attack success rates with ~73.55% fewer queries compared to the\nstate-of-the-art method for video adversarial attacks on the widely used Jester\ndataset. Overall, our algorithm exposes vulnerabilities of diverse video\nclassification models and achieves new state-of-the-art results under black-box\nsettings on two large datasets. Code is available here:\nhttps://github.com/sli057/Geo-TRAP\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC