Category
Knowledge & Reasoning
Source
Artificial Analysis
evaluation of record
Models covered
4
in our data
Data status
Live
Top score
16.7%
best on record
Top model
GLM-5.2
Zhipu AI
Updated
2026-07-30
last ingest
A benchmark designed to test LLMs on research-level physics reasoning tasks, featuring 71 composite research challenges.
Leaderboard
Top 4 of 4 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Seventy-one research-level physics challenges — the kind of problem a physics PhD student works on for days, not a textbook exercise with a published answer. Higher is better, and scores are very low: leaders reach the teens and most models score near zero. That floor is the useful signal, because it is one of the few evaluations with obvious headroom left, so movement here means more than movement on a benchmark near saturation. With only 71 problems, two lucky solves move a model several points, so treat small differences between models as noise.