Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text Classification
Neural network architectures in natural language processing often use\nattention mechanisms to produce probability distributions over input token\nrepresentations. Attention has empirically been demonstrated to improve\nperformance in various tasks, while its weights have been extensively used as\nexplanations for model predictions. Recent studies (Jain and Wallace, 2019;\nSerrano and Smith, 2019; Wiegreffe and Pinter, 2019) have showed that it cannot\ngenerally be considered as a faithful explanation (Jacovi and Goldberg, 2020)\nacross encoders and tasks. In this paper, we seek to improve the faithfulness\nof attention-based explanations for text classification. We achieve this by\nproposing a new family of Task-Scaling (TaSc) mechanisms that learn\ntask-specific non-contextualised information to scale the original attention\nweights. Evaluation tests for explanation faithfulness, show that the three\nproposed variants of TaSc improve attention-based explanations across two\nattention mechanisms, five encoders and five text classification datasets\nwithout sacrificing predictive performance. Finally, we demonstrate that TaSc\nconsistently provides more faithful attention-based explanations compared to\nthree widely-used interpretability techniques.\n
Paper
References (60)
Scroll for more · 38 remaining