Artificial Confabulation and Untrustworthiness: The Structural Limits of Transformer-Based, Next-Token AI Architectures

This is the first paper in a five-part series examining the structural limits, institutional consequences and economic mispricings of contemporary artificial intelligence. It provides the diagnostic test: why fluent performance is compatible with non-defensible output once verification is demanded. Transformer-based large language models (LLMs) are structurally incapable of verification, defined here as architectural veto authority to halt or refuse output under conditions of irreducible ignorance (V), a core component of epistemic integrity. This paper argues that the architectural absence of V, coupled with the absence of consequence-bearing Experience (E), renders LLMs fundamentally untrustworthy, regardless of scale and performance on pattern-matching benchmarks. This incapacity also prevents organisations from sustaining Sovereign Intelligence: the ability to verify, override and remain accountable for decisions made with or by AI systems. While most research treats "hallucinations" as behavioural glitches to be solved, the KEV constraint (Knowledge is massive, Experience is absent, Verification is absent) dictates that scaling only intensifies the failure mode, which is taxonomically classified as Artificially Confabulating and Untrustworthy (ACU). This confabulation is not a defect but the architecture operating as designed: fluent completion without truth. This claim concerns model-internal epistemic capacity: tool use, retrieval and external artefacts can support system-level checking, but they do not create internal verification capability and therefore cannot be treated as self-verification. The distinction matters because governance failures begin when system-level checking is mistaken for model-internal verification. The paper draws a parallel with clinical confabulation, documents the compounding failure stack (Speed-Weighted Reasoning Collapse → Architectural Confabulation → Uncalibrated Confidence) and presents systematic adversarial interrogation across frontier systems to demonstrate cross-model repeatability. Ensembling, debate and multi-agent coordination may reduce error rates in some tasks; however, agreement among generators is not evidence of truth and does not resolve the absence of verification (V = 0). The path forward is not autonomous AI but Governed Hybrid Intelligence (GHI): human judgement augmented by machine synthesis, constrained by external verification that supplies the grounding the architecture lacks. Without GHI, scaling produces more sophisticated confabulation, not more reliable intelligence. This paper establishes the diagnostic foundation. Papers Two and Three complete the architectural analysis. Papers Four and Five translate this diagnosis into institutional liability and market mispricing. For the practitioner path, see the white paper and companion instruments in the KEV (V = 0) Zenodo community. What is established here is not merely theoretical. It creates consequences.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC