AA-Omniscience: Knowledge and Hallucination Benchmark

Knowledge & ReasoningDemo Purposes Only

Category

Knowledge & Reasoning

Source

Artificial Analysis

evaluation of record

Models covered

none ingested

Data status

Demo Purposes Only

invented numbers

Top score

not measured

Top model

Updated

no ingest

Full results

artificialanalysis.ai

on the source

A benchmark measuring factual recall and hallucination across various economically relevant domains.

Leaderboard

The shape of the source chart, drawn on invented numbers. We hold no data for this evaluation.

Demo purposes only · invented numbers · not a measurement

1Claude Opus 5Adaptive Reasoning, Max EffortAnthropic
46.0%2GPT-5.6 SolmaxOpenAI
42.7%3Kimi K3Kimi
39.1%4Grok 4.5highSpaceXAI
35.8%5GLM-5.2maxZ AI
32.0%6Muse Spark 1.1xhighMeta
29.3%7Gemini 3.5 FlashhighGoogle
26.0%8Qwen3.7 MaxAlibaba
23.3%9MiniMax-M3MiniMax
20.5%10DeepSeek V4 ProReasoning, Max EffortDeepSeek
18.1%11Motif 3BetaMotif Technologies
16.4%12MiMo-V2.5-ProXiaomi
14.7%13Hy3Tencent
13.5%14Nex-N2-ProNex AGI
12.4%15InklingxhighThinking Machines
11.3%16Agnes 2.5 Pro AlphaSapiens AI
10.3%17JT-4.1 Flash 236B A21BChina Mobile
9.4%18Nemotron 3 Ultra 550B A55BReasoningNVIDIA
8.7%19Claude Opus 5Adaptive Reasoning, Xhigh EffortAnthropic
7.6%20Claude Fable 5Adaptive Reasoning, Max Effort, Opus 4.8 FallbackAnthropic
7.0%

Model names are drawn from our own roster; the scores are invented and measure nothing. See the source for the real leaderboard.

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Asks a wide set of factual questions across economically useful domains and tracks two things at once: how much the model knows, and how often it invents an answer when it does not know. Higher is better, but the design point is that a confident wrong answer costs more than admitting ignorance, so a broadly knowledgeable model can still rank badly. That makes it one of the few public numbers that measures hallucination directly rather than inferring it from something else. Recall is a moving target: the same model scores differently with search or tools enabled than it does answering from memory, so check which configuration a quoted number came from.

© 2026 NYSGPT2525 LLC