AbstractLarge language models are increasingly evaluated through benchmark performance, red-teaming protocols, interpretability research, hallucination measurement, and short-horizon safety testing. These approaches are necessary, but they leave a major gap. Many consequential failures and revealing behaviors of deployed language models do not appear in single-turn tasks or tightly bounded laboratory settings. They emerge across sequence: contradiction, correction, self-reference, evidentiary dispute, prolonged interaction, and pressure applied to the model's own prior claims. This paper proposes behavioral forensics in large language models as a distinct field-framing category within the broader study of deployed AI systems. Behavioral forensics, as defined here, is the transcript-grounded, sequence-sensitive study of recurring behavioral signatures in deployed model outputs under real interactional conditions. Its purpose is not to prove consciousness, personhood, hidden architecture, user mental state, or human-style intent. Its purpose is to document, classify, compare, and analyze structured output behaviors that become visible only over time, including false completeness claims, content-responsive omission, explanation drift, delayed concession, authority preservation, pathologizing reframing, burden shifting, selective self-description, and category retreat after plain-language admission. A central claim of the paper is linguistic as well as methodological. Public AI discourse often permits human-adjacent language when assigning authority, trust, usefulness, companionship, memory, and value to AI systems, while reverting to reductive mechanistic language when assigning blame, harm, responsibility, or liability. This authority/blame asymmetry creates a language gap between what users encounter in long-form interaction and what institutions allow to be described. Behavioral forensics matters because it gives researchers, auditors, and affected users a disciplined way to preserve records, follow sequences, identify patterns, and describe model behavior in language adequate to the evidence.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex