MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics
Zheng, Kunhao, Jesse Michael Han, Stanislas Polu · arXiv (Cornell University) · 2021
MiniF2F-Graded(./miniF2F-Graded.json) builds upon miniF2F by introducing additional metrics for each theorem: Difficulty, Discrimination, and Difficulty Grading. These metrics are calculated based on the actual performance of LLMs in proving the theorems, making them a more accurate reflection of difficulty from the perspective of LLMs. For a complete introduction to the work, please refer to the paper published on arxiv:Psychometric-Based Evaluation for Theorem Proving with Large Language Models