MiniF2F-Graded(./miniF2F-Graded.json) builds upon miniF2F by introducing additional metrics for each theorem: Difficulty, Discrimination, and Difficulty Grading. These metrics are calculated based on the actual performance of LLMs in proving the theorems, making them a more accurate reflection of difficulty from the perspective of LLMs. For a complete introduction to the work, please refer to the paper published on arxiv:Psychometric-Based Evaluation for Theorem Proving with Large Language Models
Paper
References (25)
Scroll for more · 13 remaining