A new arXiv paper studies the robustness of proof autoformalization models in Lean 4. The authors test whether models remain faithful when informal mathematical proofs are paraphrased or altered from curated ideal forms.
The work matters because autoformalization is only useful if it can handle the messiness of real mathematical writing. Models that perform well on polished datasets may fail when proofs are expressed differently.
The study adds a practical robustness lens to AI-for-mathematics benchmarks, where correctness and faithfulness are both essential.