Faithful Autoformalization via Roundtrip Verification and Repair
Researchers from arXiv have proposed a novel framework for ensuring the faithfulness of Large Language Model (LLM) autoformalization without requiring ground-truth annotations. The method employs a roundtrip verification process where natural language statements are formalized, translated back to natural language, and re-formalized. Formal tools then check for logical equivalence between the initial and secondary formalizations. If discrepancies arise, a stage-level diagnosis pinpoints the error, and a scoped repair operator attempts correction. The framework was evaluated using Claude Opus 4.6 and GPT-5.2 on statutory domains, specifically the Texas Transportation Code and Texas Parks and Wildlife Code. Results indicated that diagnosis-guided scoped repair is the most effective method, with its success relying on diagnostic reliability. Furthermore, rules failing the equivalence check exhibited significantly higher Natural Language Inference (NLI) drift compared to those passing. This study addresses critical challenges in AI reliability and legal tech by offering a robust mechanism for verifying automated logical translations.
Wire timeline
Faithful Autoformalization via Roundtrip Verification and Repair
Researchers from arXiv have proposed a novel framework for ensuring the faithfulness of Large Language Model (LLM) autoformalization without requiring ground-truth annotations. The method employs a roundtrip verification process where natural language statements are formalized, translated back to natural language, and re-formalized. Formal tools then check for logical equivalence between the initial and secondary formalizations. If discrepancies arise, a stage-level diagnosis pinpoints the error, and a scoped repair operator attempts correction. The framework was evaluated using Claude Opus 4.6 and GPT-5.2 on statutory domains, specifically the Texas Transportation Code and Texas Parks and Wildlife Code. Results indicated that diagnosis-guided scoped repair is the most effective method, with its success relying on diagnostic reliability. Furthermore, rules failing the equivalence check exhibited significantly higher Natural Language Inference (NLI) drift compared to those passing. This study addresses critical challenges in AI reliability and legal tech by offering a robust mechanism for verifying automated logical translations.
cs.AI updates on arXiv.org