Local Multi-Head Channel Self-Attention for Facial Expression Recognition

Since the Transformer architecture was introduced in 2017 there has been many\nattempts to bring the self-attention paradigm in the field of computer vision.\nIn this paper we propose a novel self-attention module that can be easily\nintegrated in virtually every convolutional neural network and that is\nspecifically designed for computer vision, the LHC: Local (multi) Head Channel\n(self-attention). LHC is based on two main ideas: first, we think that in\ncomputer vision the best way to leverage the self-attention paradigm is the\nchannel-wise application instead of the more explored spatial attention and\nthat convolution will not be replaced by attention modules like recurrent\nnetworks were in NLP; second, a local approach has the potential to better\novercome the limitations of convolution than global attention. With LHC-Net we\nmanaged to achieve a new state of the art in the famous FER2013 dataset with a\nsignificantly lower complexity and impact on the "host" architecture in terms\nof computational cost when compared with the previous SOTA.\n

Paper

References (47)

Scroll for more · 35 remaining

Similar papers

© 2026 NYSGPT2525 LLC