From Generative Capability to Assertability: Epistemic Self-Regulation Training for Large Language Models

Large language models continue to grow more capable at generation, question answering, reasoning, and tool use, yet gains in capability do not automatically translate into epistemic reliability. Their characteristic failures often lie not merely in saying something false, but in faulty epistemic regulation upstream of assertion-like output: presenting claims in a confident register when the evidence is thin, answering directly where verification is called for, offering explanations that are not faithful to the actual basis of a decision, invoking external sources without marking how they bear on the answer, and deferring to the user where the system should instead withhold or refuse. This paper does not presuppose that current LLMs are accountable asserters, testifiers, or epistemic agents in the full sense. It starts instead from the practical fact that their outputs are increasingly taken up as assertion-like contributions within human knowledge practices. We argue that the central problem in training LLMs is therefore not only to improve a model's generative capability---what it is able to produce---but to shape its assertability: the functional conditions under which an output may be responsibly presented, accepted, or withheld as an assertion. Crucially, assertability is relative to an evidence frame: the set of evidential sources admissible for the current interaction. A claim may be true, or even supported by information available elsewhere, and still fail to be assertable if it exceeds the sources authorized by the current task. To this end, we introduce epistemic self-regulation training (ESR) as a mid-level concept that unifies uncertainty discipline, mediated verification, deliberative honesty, and calibrated refusal into a single second-order framework for regulating the warrant to assert. By ``second-order,'' we mean that ESR does not primarily regulate a model's first-order capacity to generate, reason, or retrieve, but rather the conditions---of evidential state, authorized evidential sources, evidence-boundary marking, and accountable interface design---under which such capacities may be converted into assertion-like outputs. We anchor the framework in the epistemological tradition on the norms of assertion while avoiding commitment to the claim that LLMs possess human-like belief, knowledge, or responsibility. Prevailing training paradigms---pretraining, preference-based post-training, reasoning-oriented reinforcement, and deployment-time tool access---tend to treat epistemic regulation as an add-on feature or a post hoc patch rather than an independent training objective, and therefore systematically neglect to shape the model's epistemic posture. The framework requires only that LLMs, as knowledge interfaces, exhibit trainable, evaluable, and governable behavior across calibrated assertion, verification seeking, evidence-boundary marking, withholding, and refusal. We connect these dimensions to existing engineering and evaluation evidence, showing that ESR is not a purely normative manifesto but a set of training commitments that can be operationalized in objective functions, data annotation, evaluation benchmarks, and deployment interfaces. The paper contributes by reconstructing otherwise fragmented advances in calibration, retrieval, process supervision, faithfulness, and refusal as parts of a unified agenda of second-order regulation. It thereby argues that the central shift required in contemporary LLM training is not merely to enhance generative capability, but to establish assertability itself---under the evidence frame of the interaction---as an independent target of training.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC