Why do you think that? Exploring Faithful Sentence-Level Rationales Without Supervision

Evaluating the trustworthiness of a model's prediction is essential for\ndifferentiating between `right for the right reasons' and `right for the wrong\nreasons'. Identifying textual spans that determine the target label, known as\nfaithful rationales, usually relies on pipeline approaches or reinforcement\nlearning. However, such methods either require supervision and thus costly\nannotation of the rationales or employ non-differentiable models. We propose a\ndifferentiable training-framework to create models which output faithful\nrationales on a sentence level, by solely applying supervision on the target\ntask. To achieve this, our model solves the task based on each rationale\nindividually and learns to assign high scores to those which solved the task\nbest. Our evaluation on three different datasets shows competitive results\ncompared to a standard BERT blackbox while exceeding a pipeline counterpart's\nperformance in two cases. We further exploit the transparent decision-making\nprocess of these models to prefer selecting the correct rationales by applying\ndirect supervision, thereby boosting the performance on the rationale-level.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC