This paper presents BiViT, a novel hybrid vision Transformer architecture that combines the strengths of convolutional neural networks (CNNs) and attention mechanisms for image classification tasks. The proposed model introduces three key innovations: (1) FeaMixer blocks that enhance feature representation through optimized mixing operations, (2) a dual-layer routing attention mechanism that dynamically iVielects relevant regions for efficient computation, and (3) large convolutional kernels that capture long-range dependencies while maintaining computational efficiency. Extensive experiments on the Food101 dataset demonstrate that BiViT achieves superior performance compared to existing architectures, attaining 68% validation accuracy with exceptional training stability. The model reaches near-optimal performance within just 20 epochs and maintains a consistent 2-3% accuracy advantage over competing approaches. This performance improvement is attributed to BiViT’s effective hybrid design that successfully integrates Transformer and CNN components, optimized feature mixing mechanisms, and enhanced regularization techniques. The results suggest that BiViT represents a significant advancement in efficient and effective visual representation learning.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex