Interactive Critique-Revision Training for Reliable Structured LLM Generation
Researchers have introduced DPA-GRPO (Dual Paired-Action Group-Relative Policy Optimization), a novel training method designed to enhance the reliability of Large Language Models (LLMs) in structured decision-making tasks such as form filling and compliance checking. Addressing the limitations of existing refinement methods that often create second-order assurance problems, this approach utilizes a two-player generator-verifier game. The generator proposes outputs and revises them when challenged, while the verifier issues Safety Assurance Cases (SACs) or remains silent. These interactions create paired counterfactual action groups for role-specific policy updates. Theoretical analysis demonstrates that the method tracks stable limit points under stochastic approximation. Experimental results on the TaxCalcBench TY24 dataset indicate that DPA-GRPO significantly improves structured decision accuracy compared to zero-shot generation and generator-only reinforcement learning baselines across Qwen3-4B and Qwen3-8B models. The training process increases correct silent acceptance, reduces missed errors, and improves calibrated revision behavior, offering substantial gains for both generator and verifier components in ensuring locally correct and globally consistent LLM outputs.
Wire timeline
Interactive Critique-Revision Training for Reliable Structured LLM Generation
Researchers have introduced DPA-GRPO (Dual Paired-Action Group-Relative Policy Optimization), a novel training method designed to enhance the reliability of Large Language Models (LLMs) in structured decision-making tasks such as form filling and compliance checking. Addressing the limitations of existing refinement methods that often create second-order assurance problems, this approach utilizes a two-player generator-verifier game. The generator proposes outputs and revises them when challenged, while the verifier issues Safety Assurance Cases (SACs) or remains silent. These interactions create paired counterfactual action groups for role-specific policy updates. Theoretical analysis demonstrates that the method tracks stable limit points under stochastic approximation. Experimental results on the TaxCalcBench TY24 dataset indicate that DPA-GRPO significantly improves structured decision accuracy compared to zero-shot generation and generator-only reinforcement learning baselines across Qwen3-4B and Qwen3-8B models. The training process increases correct silent acceptance, reduces missed errors, and improves calibrated revision behavior, offering substantial gains for both generator and verifier components in ensuring locally correct and globally consistent LLM outputs.
cs.AI updates on arXiv.org