HiProtoAD: Feature Reconstruction Guided by Cluster-Separated Hierarchical Prototypes for Multi-Class Unsupervised Anomaly Detection
Prevailing Transformer-based feature reconstruction methods for multi-class unsupervised visual anomaly detection suffer from identity mapping, while the underlying causes remain insufficiently understood. Through a systematic analysis, we reveal that identity mapping is closely related to anomaly injection arising from cross-category information interaction and model architectures. To address this issue, we propose HiProtoAD, a novel framework that guides feature reconstruction via cluster-separated hierarchical prototype learning. Built upon a Vision Transformer–based hierarchical reconstruction architecture, HiProtoAD incorporates three key components: 1) a Cluster-Separated Learnable Token Selector (CSLTS) to isolate category-specific normal prototype extraction and mitigate cross-category anomaly injection; 2) a Sparse Cross-Attention (SCA) module to suppress low-relevance anomalous information; and 3) a Gaussian-Masked Cross-Attention (GMCA) module to enhance normal prototype-guided feature reconstruction. Extensive experiments on benchmark datasets, including MVTec-AD, VisA, BTAD, and MPDD, demonstrate that HiProtoAD achieves competitive performance compared with state-of-the-art methods, providing an effective solution for complex multi-class unsupervised visual anomaly detection.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex