Human Parity on CommonsenseQA: Augmenting Self-Attention with External Attention

Most of today's AI systems focus on using self-attention mechanisms and\ntransformer architectures on large amounts of diverse data to achieve\nimpressive performance gains. In this paper, we propose to augment the\ntransformer architecture with an external attention mechanism to bring external\nknowledge and context to bear. By integrating external information into the\nprediction process, we hope to reduce the need for ever-larger models and\nincrease the democratization of AI systems. We find that the proposed external\nattention mechanism can significantly improve the performance of existing AI\nsystems, allowing practitioners to easily customize foundation AI models to\nmany diverse downstream applications. In particular, we focus on the task of\nCommonsense Reasoning, demonstrating that the proposed external attention\nmechanism can augment existing transformer models and significantly improve the\nmodel's reasoning capabilities. The proposed system, Knowledgeable External\nAttention for commonsense Reasoning (KEAR), reaches human parity on the open\nCommonsenseQA research benchmark with an accuracy of 89.4\\% in comparison to\nthe human accuracy of 88.9\\%.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC