REFUTE is an open benchmark for scientific critique honesty and epistemic calibration in language models. Homepage: https://bgpt.pro/refute Preprint DOI: https://doi.org/10.5281/zenodo.21542238 Code: https://github.com/connerlambden/refute-inspect bio.tools: https://bio.tools/bgpt-refute OpenML: https://www.openml.org/d/47266 and https://www.openml.org/d/47268 Wikidata: https://www.wikidata.org/wiki/Q140690703 Measures whether models correctly critique scientific claims and whether their confidence is calibrated.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex