MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics

MiniF2F-Graded(./miniF2F-Graded.json) builds upon miniF2F by introducing additional metrics for each theorem: Difficulty, Discrimination, and Difficulty Grading. These metrics are calculated based on the actual performance of LLMs in proving the theorems, making them a more accurate reflection of difficulty from the perspective of LLMs. For a complete introduction to the work, please refer to the paper published on arxiv:Psychometric-Based Evaluation for Theorem Proving with Large Language Models

Paper

References (25)

Scroll for more · 13 remaining

Similar papers

© 2026 NYSGPT2525 LLC