Neural Natural Language Inference Models Partially Embed Theories of Lexical Entailment and Negation

We address whether neural models for Natural Language Inference (NLI) can\nlearn the compositional interactions between lexical entailment and negation,\nusing four methods: the behavioral evaluation methods of (1) challenge test\nsets and (2) systematic generalization tasks, and the structural evaluation\nmethods of (3) probes and (4) interventions. To facilitate this holistic\nevaluation, we present Monotonicity NLI (MoNLI), a new naturalistic dataset\nfocused on lexical entailment and negation. In our behavioral evaluations, we\nfind that models trained on general-purpose NLI datasets fail systematically on\nMoNLI examples containing negation, but that MoNLI fine-tuning addresses this\nfailure. In our structural evaluations, we look for evidence that our\ntop-performing BERT-based model has learned to implement the monotonicity\nalgorithm behind MoNLI. Probes yield evidence consistent with this conclusion,\nand our intervention experiments bolster this, showing that the causal dynamics\nof the model mirror the causal dynamics of this algorithm on subsets of MoNLI.\nThis suggests that the BERT model at least partially embeds a theory of lexical\nentailment and negation at an algorithmic level.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC