Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning
Researchers from arXiv have proposed a novel method called failure-prefix conditioning to address a emerging bottleneck in training large language models (LLMs) using Reinforcement Learning with Verifiable Rewards (RLVR). As LLMs improve, many training problems become saturated, meaning the models answer correctly nearly every time, providing little useful learning signal. Collecting harder problems is costly and difficult. The proposed method shifts exploration toward failure-prone reasoning states by conditioning on prefixes of rare incorrect trajectories. This approach helps models recover from misleading early reasoning steps. Results indicate that failure-prefix conditioning consistently improves performance where standard RLVR stalls, achieving gains comparable to training on new medium-difficulty problems. The study also notes a mild trade-off in adhering to correct early reasoning but highlights improved robustness against misleading prefixes. An iterative approach refreshing these prefixes during training unlocks further gains after performance plateaus. The findings suggest that saturated problems still hold valuable learning signals that can be effectively utilized through this technique, offering a cost-effective alternative to sourcing new datasets for enhancing LLM reasoning capabilities.
Wire timeline
Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning
Researchers from arXiv have proposed a novel method called failure-prefix conditioning to address a emerging bottleneck in training large language models (LLMs) using Reinforcement Learning with Verifiable Rewards (RLVR). As LLMs improve, many training problems become saturated, meaning the models answer correctly nearly every time, providing little useful learning signal. Collecting harder problems is costly and difficult. The proposed method shifts exploration toward failure-prone reasoning states by conditioning on prefixes of rare incorrect trajectories. This approach helps models recover from misleading early reasoning steps. Results indicate that failure-prefix conditioning consistently improves performance where standard RLVR stalls, achieving gains comparable to training on new medium-difficulty problems. The study also notes a mild trade-off in adhering to correct early reasoning but highlights improved robustness against misleading prefixes. An iterative approach refreshing these prefixes during training unlocks further gains after performance plateaus. The findings suggest that saturated problems still hold valuable learning signals that can be effectively utilized through this technique, offering a cost-effective alternative to sourcing new datasets for enhancing LLM reasoning capabilities.
cs.AI updates on arXiv.org