What Changed

Large language models (LLMs) have demonstrated proficiency in high-school and olympiad-style mathematics. However, their performance in advanced mathematics has remained less understood due to limitations in existing benchmarks. These prior benchmarks often lacked sufficient disciplinary scope and relied on coarse evaluation methods, such as final-answer correctness, which failed to adequately assess the validity of the reasoning process itself.

To address this gap, a new benchmark suite, AdvancedMathBench, has been introduced. This suite is specifically designed to evaluate LLMs' capabilities in advanced mathematical reasoning, focusing on both proof generation and verification. AdvancedMathBench aims to provide a more comprehensive and granular assessment of LLM performance in complex mathematical tasks, moving beyond simpler problem sets to tackle challenges at the undergraduate and doctoral qualifying-exam levels.

Technical Details

AdvancedMathBench comprises two core components: ProverBench and VerifierBench.