NLI Data Sanity Check: Assessing the Effect of Data Corruption on Model Performance

Pre-trained neural language models give high performance on natural language\ninference (NLI) tasks. But whether they actually understand the meaning of the\nprocessed sequences remains unclear. We propose a new diagnostics test suite\nwhich allows to assess whether a dataset constitutes a good testbed for\nevaluating the models' meaning understanding capabilities. We specifically\napply controlled corruption transformations to widely used benchmarks (MNLI and\nANLI), which involve removing entire word classes and often lead to\nnon-sensical sentence pairs. If model accuracy on the corrupted data remains\nhigh, then the dataset is likely to contain statistical biases and artefacts\nthat guide prediction. Inversely, a large decrease in model accuracy indicates\nthat the original dataset provides a proper challenge to the models' reasoning\ncapabilities. Hence, our proposed controls can serve as a crash test for\ndeveloping high quality data for NLI tasks.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC