Exploring the impact of zero-cost proxies for Hybrid Vision Transformers

Image classification plays a crucial role in numerous modern applications, including visual inspection, autonomous vehicles, demographics, surveillance, healthcare, and human-computer interaction. This encompasses various computer vision tasks such as object recognition, optical character recognition, facial recognition, and more. Deep Neural Networks (DNN) achieved state-of-the-art performance in many of these tasks. In particular, Vision Transformers (ViT) have shown outstanding results. However, their manual design may not be practical in resource-constrained scenarios. To address this issue, state-of-the-art Neural Architecture Search (NAS) automate DNN design by employing intelligent search strategies combined with proxies to estimate models’ performance without training. In this study, we evaluate zero-cost proxies in the context of NAS for hybrid ViT. Experiments were conducted on image classification datasets for different tasks. Our findings support the notion that convolutions and multi-head self-attention exhibit distinct properties, and therefore, different criteria are required to predict their performances accurately.

Paper

Full text

PDF

Exploring the impact of zero-cost proxies for Hybrid Vision Transformers

Semantic Scholar · Computer Science · 2024

Abstract

Image classification plays a crucial role in numerous modern applications, including visual inspection, autonomous vehicles, demographics, surveillance, healthcare, and human-computer interaction. This encompasses various computer vision tasks such as object recognition, optical character recognition, facial recognition, and more. Deep Neural Networks (DNN) achieved state-of-the-art performance in many of these tasks. In particular, Vision Transformers (ViT) have shown outstanding results. However, their manual design may not be practical in resource-constrained scenarios. To address this issue, state-of-the-art Neural Architecture Search (NAS) automate DNN design by employing intelligent search strategies combined with proxies to estimate models’ performance without training. In this study, we evaluate zero-cost proxies in the context of NAS for hybrid ViT. Experiments were conducted on image classification datasets for different tasks. Our findings support the notion that convolutions and multi-head self-attention exhibit distinct properties, and therefore, different criteria are required to predict their performances accurately.

Similar papers

© 2026 NYSGPT2525 LLC