ANCORA: Learning to Question via Manifold-Anchored Self-Play for Verifiable Reasoning
Researchers have introduced ANCORA, a novel machine learning framework that shifts the paradigm from answering fixed prompts to generating verifiable problems through open-ended curriculum self-play. The system employs a unified policy alternating between a Proposer, which synthesizes novel specifications, and a Solver, which produces verified solutions. This process operates without human-annotated solutions, relying on three key stabilizing mechanisms: two-level group-relative updates, iterative self-distilled supervised fine-tuning, and a UCB-guided Curriculum DAG. These components prevent proposer collapse and enable the model to bootstrap a verifiable curriculum from scratch. When instantiated in Verus, ANCORA significantly improved performance on the Dafny2Verus task, raising pass@1 accuracy from a 26.6% baseline to 81.5% in test-time training. It outperformed existing PSV self-play methods by 15.8 points. Additionally, the model demonstrated strong transfer capabilities, achieving notable pass@1 scores on held-out MBPP and HumanEval benchmarks. This development marks a significant advancement in autonomous AI reasoning and self-improvement capabilities.
Wire timeline
ANCORA: Learning to Question via Manifold-Anchored Self-Play for Verifiable Reasoning
Researchers have introduced ANCORA, a novel machine learning framework that shifts the paradigm from answering fixed prompts to generating verifiable problems through open-ended curriculum self-play. The system employs a unified policy alternating between a Proposer, which synthesizes novel specifications, and a Solver, which produces verified solutions. This process operates without human-annotated solutions, relying on three key stabilizing mechanisms: two-level group-relative updates, iterative self-distilled supervised fine-tuning, and a UCB-guided Curriculum DAG. These components prevent proposer collapse and enable the model to bootstrap a verifiable curriculum from scratch. When instantiated in Verus, ANCORA significantly improved performance on the Dafny2Verus task, raising pass@1 accuracy from a 26.6% baseline to 81.5% in test-time training. It outperformed existing PSV self-play methods by 15.8 points. Additionally, the model demonstrated strong transfer capabilities, achieving notable pass@1 scores on held-out MBPP and HumanEval benchmarks. This development marks a significant advancement in autonomous AI reasoning and self-improvement capabilities.
cs.AI updates on arXiv.org