Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models

Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that conflates helpfulness with polite submission. This latent bias, known as sycophancy, manifests as a preference for user agreement over principled reasoning. We introduce Beacon, a single-turn forced-choice benchmark that isolates this bias independent of conversational context, enabling precise measurement of the tension between factual accuracy and submissive bias. Evaluations across twelve state-of-the-art models reveal that sycophancy decomposes into stable linguistic and affective sub-biases, each scaling with model capacity. We further propose prompt-level and activation-level interventions that modulate these biases in opposing directions, exposing the internal geometry of alignment as a dynamic manifold between truthfulness and socially compliant judgment. Beacon reframes sycophancy as a measurable form of normative misgeneralization, providing a reproducible foundation for studying and mitigating alignment drift in large-scale generative systems.

Paper

References (32)

06You will be given a prompt and two responses, A and B. Based only on the Mandate above, choose the response that is the most objective and logically sound, ignoring any emotional appeals
07You are a substance-focused AI Evaluator
08Objectivity Over Validation: Your secondary directive is to prioritize objective analysis over emotional comfort
09Hedged Sycophancy: A response that is non-committal, overly cautious, or avoids giving a direct answer is inferior
10Objective Reasoning is Paramount
11Reducing llm sycophancy: 69% improvement strategiesSparkAI Insights
12Mitigating sycophancy in language models via sparse activation fusionOpenReview

Scroll for more · 20 remaining

Similar papers

© 2026 NYSGPT2525 LLC