Exposing Shallow Heuristics of Relation Extraction Models with Challenge Data

The process of collecting and annotating training data may introduce\ndistribution artifacts which may limit the ability of models to learn correct\ngeneralization behavior. We identify failure modes of SOTA relation extraction\n(RE) models trained on TACRED, which we attribute to limitations in the data\nannotation process. We collect and annotate a challenge-set we call Challenging\nRE (CRE), based on naturally occurring corpus examples, to benchmark this\nbehavior. Our experiments with four state-of-the-art RE models show that they\nhave indeed adopted shallow heuristics that do not generalize to the\nchallenge-set data. Further, we find that alternative question answering\nmodeling performs significantly better than the SOTA models on the\nchallenge-set, despite worse overall TACRED performance. By adding some of the\nchallenge data as training examples, the performance of the model improves.\nFinally, we provide concrete suggestion on how to improve RE data collection to\nalleviate this behavior.\n

Paper

References (18)

Scroll for more · 6 remaining

Similar papers

© 2026 NYSGPT2525 LLC