We introduce exclusive self attention (XSA), a simple modification of self attention (SA) that improves Transformer's sequence modeling performance. The key idea is to constrain attention to capture only information orthogonal to the token's own value vector (thus excluding information of self position), encouraging better context modeling. Evaluated on the standard language modeling task, XSA consistently outperforms SA across model sizes up to 2.7B parameters and shows increasingly larger gains as sequence length grows.
Paper
References (15)
07Layer NormalizationJimmy Ba, J. Kiros, Geoffrey E. Hinton2016 · arXiv.org · 13k citations In Library
11language models with attention sinks. arXiv preprint arXiv:2309.174532019 · Proceedings of the 57th annual meeting of the association for computational linguistics
12Social iqa: Common-sense reasoning about social interactions2019 · Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP)
Scroll for more · 3 remaining