ITPEval: Benchmarking Automated Translation Across Major Interactive Theorem Provers
A new arXiv preprint introduces ITPEval, a benchmark designed to evaluate automated translation of formal proofs between four widely used interactive theorem provers: Lean 4, Rocq, Isabelle, and HOL Light. The benchmark includes over 1,500 source files and nearly 7,000 theorems, and assesses both statement and proof translation using several large language models. Results show that proof translation remains challenging, with a maximum pass@1 rate of 10.5%, and that mismatches between libraries are a major obstacle.
Why it matters: This work provides the first large-scale, multi-system benchmark for proof translation, highlighting key challenges for interoperability and data sharing in formal mathematics and AI-driven theorem proving.
Full story at: arXiv AI/ML ↗