SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology
Researchers have introduced Structure-Aware Reinforcement Learning (SARL), a novel label-free framework designed to enhance large reasoning models in open-ended domains where correctness is ambiguous. Unlike traditional Reinforcement Learning with Verifiable Rewards (RLVR), which focuses on final outcomes, SARL rewards the topology of reasoning maps constructed from intermediate thinking steps. This approach shifts supervision from the destination to the path, encouraging locally coherent and globally efficient reasoning trajectories. Experimental results demonstrate that SARL outperforms prior label-free baselines and even exceeds methods using ground truth supervision. On verifiable math tasks, it achieved average gains of +9.1% under PPO and +11.6% under GRPO, with significant improvements on AIME25. For non-verifiable open-ended tasks on WildBench, SARL showed average gains of +34.6% under PPO and +30.4% under GRPO, surpassing DPO. The framework also exhibits lower KL divergence and higher policy entropy, indicating more stable and exploratory training dynamics. The code and data for this study are publicly available, marking a significant advancement in applying reinforcement learning to complex, unverified reasoning tasks without relying on explicit labels.
Wire timeline
SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology
Researchers have introduced Structure-Aware Reinforcement Learning (SARL), a novel label-free framework designed to enhance large reasoning models in open-ended domains where correctness is ambiguous. Unlike traditional Reinforcement Learning with Verifiable Rewards (RLVR), which focuses on final outcomes, SARL rewards the topology of reasoning maps constructed from intermediate thinking steps. This approach shifts supervision from the destination to the path, encouraging locally coherent and globally efficient reasoning trajectories. Experimental results demonstrate that SARL outperforms prior label-free baselines and even exceeds methods using ground truth supervision. On verifiable math tasks, it achieved average gains of +9.1% under PPO and +11.6% under GRPO, with significant improvements on AIME25. For non-verifiable open-ended tasks on WildBench, SARL showed average gains of +34.6% under PPO and +30.4% under GRPO, surpassing DPO. The framework also exhibits lower KL divergence and higher policy entropy, indicating more stable and exploratory training dynamics. The code and data for this study are publicly available, marking a significant advancement in applying reinforcement learning to complex, unverified reasoning tasks without relying on explicit labels.
cs.AI updates on arXiv.org