Accuracy Without Stability Is Not Intelligence: Behavioral Evaluation as the Missing Dimension in LLM Benchmarks
This paper demonstrates that output correctness is a necessary but insufficient indicator of large language model reliability. Through 2,040 instrumented observations across two production models and a 27-model behavioral benchmark, we identify failure modes invisible to accuracy-based evaluation: models producing correct outputs while operating in unstable internal states, reasoning that changes under minimal perturbation, and self-contradiction under verification. We propose three observable behavioral tests — Consistency, Continuity, and Self-Agreement — requiring no access to model internals. Companion paper to the Cognitive Entropy Collapse Battery (CECB) Kaggle competition submission. Published figures embedded in HTML.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex