CritPt Benchmark Leaderboard

Knowledge & ReasoningLive

Category

Knowledge & Reasoning

Source

Artificial Analysis

evaluation of record

Models covered

4

in our data

Data status

Live

Top score

16.7%

best on record

Top model

GLM-5.2

Zhipu AI

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A benchmark designed to test LLMs on research-level physics reasoning tasks, featuring 71 composite research challenges.

Leaderboard

Top 4 of 4 models we hold a score for.

1GLM-5.2Zhipu AI
16.7%2Qwen3.7 MaxAlibaba Cloud / Qwen Team
11.4%3Qwen3.7-PlusAlibaba Cloud / Qwen Team
6.0%4Nemotron 3 Ultra550B A55BNVIDIA
3.1%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Seventy-one research-level physics challenges — the kind of problem a physics PhD student works on for days, not a textbook exercise with a published answer. Higher is better, and scores are very low: leaders reach the teens and most models score near zero. That floor is the useful signal, because it is one of the few evaluations with obvious headroom left, so movement here means more than movement on a benchmark near saturation. With only 71 problems, two lucky solves move a model several points, so treat small differences between models as noise.

© 2026 NYSGPT2525 LLC