REFUTE: Scientific Critique and Epistemic Calibration Benchmark

REFUTE is an open benchmark for scientific critique honesty and epistemic calibration in language models. Homepage: https://bgpt.pro/refute Preprint DOI: https://doi.org/10.5281/zenodo.21542238 Code: https://github.com/connerlambden/refute-inspect bio.tools: https://bio.tools/bgpt-refute OpenML: https://www.openml.org/d/47266 and https://www.openml.org/d/47268 Wikidata: https://www.wikidata.org/wiki/Q140690703 Measures whether models correctly critique scientific claims and whether their confidence is calibrated.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC