Large language models deployed in production environments exhibit behavioral drift gradual changes in output characteristics that degrade reliability over time. Despite exten- sive research on static LLM evaluation, the temporal dynamics of model behavior remain poorly understood. We introduce DriftEval, a continuous evaluation framework that detects and quanties multi-dimensional drift in LLM outputs across semantic, stylistic, behavioral, and factual dimensions. Our framework employs embedding-based divergence metrics, statistical distribution tests, and task-specic consistency measures to identify drift before it impacts downstream applications. Experiments tracking GPT-3.5, GPT-4, and Claude-2 over a 6-month period reveal systematic drift patterns: semantic drift coecients ranging from 0.12 to 0.34, stylistic variation increases of 23% on average, and behavioral pattern shifts aecting 17% of response categories. We demonstrate that DriftEval de- tects statistically signicant drift (p < 0.01) an average of 2.3 weeks before user-reported quality degradation. Our framework enables proactive monitoring, automated alerting, and temporal analysis for production LLM deployments.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex