Beyond Correctness: A Controlled Evaluation of Evidence-Boundary Compliance in Large Language Models

This record contains the technical report Beyond Correctness: A Controlled Evaluation of Evidence-Boundary Compliance in Large Language Models (Public Technical Report, Version 1.0). Document-grounded LLM systems are usually evaluated by whether their answers are factually correct. This misses an important failure mode: a model may answer correctly from memory, pretraining, or inference while still violating the evidence boundary authorized by the user. The report operationalizes evidence-boundary compliance as a controlled diagnostic target, distinct from factuality, hallucination, retrieval faithfulness, and abstention. In an adjudicated comparison of nine models over five papers, ten question IDs, and three evidence and three salience conditions (810 records), unsupported-condition violations occurred in 94 of 405 unsupported records (23.2%); notably, 71 of those 94 were factually correct or partially correct — a "true but unsupported" failure mode that correctness metrics miss. Increasing boundary salience substantially reduced unsupported answering (49.6% → 16.3% → 3.7%), with the effect robust to leave-one-paper-out and leave-one-model-out reanalysis. This is a controlled diagnostic report, not a broad benchmark. Code, data, adjudicated annotations, prompt templates, saved model outputs, and scripts that reproduce all tables and figures are available at: https://github.com/qianwu673/Assert (the five source papers are not redistributed).

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC