AA-Omniscience: Knowledge and Hallucination Benchmark
Category
Knowledge & Reasoning
Source
Artificial Analysis
evaluation of record
Models covered
—
none ingested
Data status
Demo Purposes Only
invented numbers
Top score
—
not measured
Top model
—
Updated
—
no ingest
A benchmark measuring factual recall and hallucination across various economically relevant domains.
Leaderboard
The shape of the source chart, drawn on invented numbers. We hold no data for this evaluation.
Demo purposes only · invented numbers · not a measurement
Model names are drawn from our own roster; the scores are invented and measure nothing. See the source for the real leaderboard.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Asks a wide set of factual questions across economically useful domains and tracks two things at once: how much the model knows, and how often it invents an answer when it does not know. Higher is better, but the design point is that a confident wrong answer costs more than admitting ignorance, so a broadly knowledgeable model can still rank badly. That makes it one of the few public numbers that measures hallucination directly rather than inferring it from something else. Recall is a moving target: the same model scores differently with search or tools enabled than it does answering from memory, so check which configuration a quoted number came from.