Capturing AI's Attention: Physics of Repetition, Hallucination, Bias and Beyond

We derive a first-principles physics theory of the AI engine at the heart of LLMs' 'magic' (e.g. ChatGPT, Claude): the basic Attention head. The theory allows a quantitative analysis of outstanding AI challenges such as output repetition, hallucination and harmful content, and bias (e.g. from training and fine-tuning). Its predictions are consistent with large-scale LLM outputs. Its 2-body form suggests why LLMs work so well, but hints that a generalized 3-body Attention would make such AI work even better. Its similarity to a spin-bath means that existing Physics expertise could immediately be harnessed to help Society ensure AI is trustworthy and resilient to manipulation.

Paper

References (15)

09Attention in nat-ural language processing2021 · IEEE Transactions on Neural Networks and Learning Systems
10A mechanistic interpretability analysis of grokking
11Doge will use AI to assess the responses of federal workers who were told to justify their jobs via email
12The AI relationship revolution is already hereMIT Technology Review

Scroll for more · 3 remaining

Similar papers

© 2026 NYSGPT2525 LLC