Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs
Researchers have introduced Forge and OPT-BENCH, a new framework designed to enhance Large Language Models (LLMs) in solving NP-hard optimization problems. Unlike existing benchmarks that focus solely on correctness, this approach utilizes quality-aware Reinforcement Learning with Verifiable Rewards (RLVR) to evaluate optimality. The framework includes scalable training infrastructure, quality verifiers, and a rigorous benchmark of 1,000 instances across ten tasks. Experimental results demonstrate that training the Qwen2.5-7B-Instruct-1M model on 15,000 examples yields a 93.1% Success Rate and 46.6% Quality Ratio, significantly outperforming GPT-4o, which achieved 29.6% and 14.6% respectively. Furthermore, the study reveals that quality-aware rewards improve solution quality by 28.8% compared to binary rewards. The training also shows positive transfer effects to other domains, including mathematics, logic, knowledge, and instruction following. The findings suggest that task diversity is more critical than data quantity for generalization in complex reasoning tasks, offering valuable insights for scaling RLVR in AI development.
Wire timeline
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs
Researchers have introduced Forge and OPT-BENCH, a new framework designed to enhance Large Language Models (LLMs) in solving NP-hard optimization problems. Unlike existing benchmarks that focus solely on correctness, this approach utilizes quality-aware Reinforcement Learning with Verifiable Rewards (RLVR) to evaluate optimality. The framework includes scalable training infrastructure, quality verifiers, and a rigorous benchmark of 1,000 instances across ten tasks. Experimental results demonstrate that training the Qwen2.5-7B-Instruct-1M model on 15,000 examples yields a 93.1% Success Rate and 46.6% Quality Ratio, significantly outperforming GPT-4o, which achieved 29.6% and 14.6% respectively. Furthermore, the study reveals that quality-aware rewards improve solution quality by 28.8% compared to binary rewards. The training also shows positive transfer effects to other domains, including mathematics, logic, knowledge, and instruction following. The findings suggest that task diversity is more critical than data quantity for generalization in complex reasoning tasks, offering valuable insights for scaling RLVR in AI development.
cs.AI updates on arXiv.org