Mathematics benchmarks and reproducibility efforts expand
New benchmarks (MA-ProofBench, ComBench, DynaMath 2026, MIRA-Math) reveal gaps. FrontierMath expands to ~50 unsolved problems; AI solves 3 (6%). Leanstral 1.5 achieves 99.8% on miniF2F. Critical audits of Lean benchmarks show dataset defects. New self-distillation method (PS-OPSD) improves math reasoning.
Sources (2)
Updated Aug 7, 2026