Three Dimension Dynamic Attention Meets Self-Attention

We propose a novel attention mechanism called Three Dimension Dynamic Attention, which combines global attention to efficiently handle various classification and object detection tasks. It consists of three key components. Firstly, it introduces the Four-Head-Three-Path Module, which not only captures multi-scale contextual features but also enhances local features while significantly expanding the receptive field. Secondly, the Information Interconnection Module not only integrates information between different heads but also aggregates information among the three paths. Importantly, it is highly efficient and lightweight. Lastly, the Feature Alignment Module maps the features obtained from the second step into three-dimensional space, providing 3D dynamic attention for the feature map. The effective combination of these three modules enhances feature extraction capabilities and feature long-range transmission capabilities. To better utilize the advantages of local modeling and holistic representation, an improved hybrid network is proposed. It adaptively allocates the most reasonable computing resources based on the complexity of the modules to optimize computing resource allocation. Experimental results demonstrate that the network achieves an 84.2% Top-1 accuracy on the ImageNet datasets. For object detection with COCO datasets, it surpasses the Focal-T 1.9mAP at approximately equivalent computational costs and parameters

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC