Multi-layer Attention Aggregation in Deep Neural Network

Convolutional neural networks have achieved significant successes in image classification recently due to its high capacity in learning discriminative features. In this work, we propose Multi-layer attention aggregation(MAA) model, a convolutional architecture using attention mechanism and global aggregation module iteratively. By merging attention-aware features from every convolutional stage, MAA improves the performance of image classification. Specifically, the proposed MAA model can be applied to state-of-the-art convolutional architectures, such as ResNet, and improve its performance by increasing 30% computational cost. Furthermore, we also employ ArcFace loss in the training process to improve the performance of image classification. Applying the proposed method on ResNet, our MAA model achieves higher image classification performance including on standard benchmarks of Google-Landmarks dataset, CIFAR-10 and CIFAR-100 dataset. Note that, our method achieves 0.68% top-1 accuracy improvement on Google-Landmarks dataset, 2.27% top-1 accuracy improvement on CIFAR-100 and 1.14% top-1 accuracy improvement on CIFAR-10.

Paper

Full text

PDF

Multi-layer Attention Aggregation in Deep Neural Network

Semantic Scholar · Computer Science · 2019

Abstract

Convolutional neural networks have achieved significant successes in image classification recently due to its high capacity in learning discriminative features. In this work, we propose Multi-layer attention aggregation(MAA) model, a convolutional architecture using attention mechanism and global aggregation module iteratively. By merging attention-aware features from every convolutional stage, MAA improves the performance of image classification. Specifically, the proposed MAA model can be applied to state-of-the-art convolutional architectures, such as ResNet, and improve its performance by increasing 30% computational cost. Furthermore, we also employ ArcFace loss in the training process to improve the performance of image classification. Applying the proposed method on ResNet, our MAA model achieves higher image classification performance including on standard benchmarks of Google-Landmarks dataset, CIFAR-10 and CIFAR-100 dataset. Note that, our method achieves 0.68% top-1 accuracy improvement on Google-Landmarks dataset, 2.27% top-1 accuracy improvement on CIFAR-100 and 1.14% top-1 accuracy improvement on CIFAR-10.

Similar papers

© 2026 NYSGPT2525 LLC