PETAH: Parameter Efficient Task Adaptation for Hybrid Transformers in a resource-limited Context
Transformers have revolutionized natural language processing (NLP) and are increasingly influential in computer vision tasks. Despite their strong performance and multitasking capabilities, transformers' high computational demands limit their applicability in resource-constrained environments, where convolutional or hybrid models (combining convolution and attention layers) often excel, particularly in the sub-100M parameter range. While parameterefficient task adaptation techniques have been successful in NLP, they have not been widely adopted for hybrid transformers in vision tasks. In this work, we introduce PETAH (Parameter Efficient Task Adaptation for Hybrid Transformers), a novel framework for efficiently adapting hybrid transformers to new tasks. We further combine PETAH with pruning to create high-performing and storage-efficient models suitable for multi-tasking. Our extensive evaluations on classification and other vision tasks demonstrate that PETAH-adapted hybrid models outperform established task-adaptation techniques for Vision Transformers (ViTs), requiring fewer parameters and achieving greater efficiency on mobile hardware.
Paper
References (84)
Scroll for more · 38 remaining