Viewpoint invariance remains challenging for visual recognition in the 3D world, as altering the viewing directions can significantly impact predictions for the same object. While substantial efforts have been dedicated to making neural networks invariant to 2D image translations and rotations, viewpoint invariance is rarely investigated. Motivated by the success of adversarial training in enhancing model robustness, we propose Viewpoint-Invariant Adversarial Training (VIAT) to improve the viewpoint robustness of the vision model. Regarding viewpoint change as an attack, we formulate VIAT as a minimax optimization problem, where the inner maximization characterizes diverse adversarial viewpoints by learning a Gaussian mixture distribution based on the proposed attack method GMVFool. The outer minimization seeks a viewpoint-invariant model by minimizing the expected loss over the worst-case viewpoint distribution. To accommodate models of varying parameter scales, we introduce two training strategies: VIAT-FP (Full Parameter Fine-tuning) and VIAT-PEIT (Parameter-Efficient Instruction-Tuning), serving as implementation variants of VIAT. For evaluation, we contribute ImageNet-V+, a large-scale dataset for empirical benchmarking of viewpoint robustness, covering various evaluation protocols, including image recognition, visual question answering, and visual entailment. Additionally, we introduce ViewRS, a certified viewpoint robustness method that provides theoretically guaranteed metrics for evaluating viewpoint invariance. Experimental results demonstrate that our methods significantly improve the viewpoint robustness of vision models with various architectures ranging from traditional task-specific models (e.g., CNN and ViT) to advanced multimodal large language models.
Paper
References (77)
Scroll for more · 38 remaining