Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
Researchers have introduced Self-ReSET, a novel pure reinforcement learning framework designed to enhance the safety robustness of Large Reasoning Models (LRMs). While LRMs exhibit strong self-correction capabilities in general domains, they often fail to recover from unsafe reasoning paths when subjected to adversarial attacks. Current alignment methods rely on static expert data, which creates a distribution mismatch with the model's dynamic reasoning traces, limiting their effectiveness. Self-ReSET addresses this by enabling models to learn from their own safety error trajectories, using these failures as initial states for reinforcement learning. This approach allows LRMs to develop intrinsic self-recovery patterns, effectively identifying and correcting unsafe intermediate states to return to benign reasoning paths. Extensive experiments demonstrate that Self-ReSET significantly improves resistance to out-of-distribution jailbreak prompts and adversarial attacks while maintaining general utility and ensuring efficient data usage. The study highlights a shift towards dynamic, on-policy training methods for better AI alignment. The associated code and data have been made publicly available to support further research in this area.
Wire timeline
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
Researchers have introduced Self-ReSET, a novel pure reinforcement learning framework designed to enhance the safety robustness of Large Reasoning Models (LRMs). While LRMs exhibit strong self-correction capabilities in general domains, they often fail to recover from unsafe reasoning paths when subjected to adversarial attacks. Current alignment methods rely on static expert data, which creates a distribution mismatch with the model's dynamic reasoning traces, limiting their effectiveness. Self-ReSET addresses this by enabling models to learn from their own safety error trajectories, using these failures as initial states for reinforcement learning. This approach allows LRMs to develop intrinsic self-recovery patterns, effectively identifying and correcting unsafe intermediate states to return to benign reasoning paths. Extensive experiments demonstrate that Self-ReSET significantly improves resistance to out-of-distribution jailbreak prompts and adversarial attacks while maintaining general utility and ensuring efficient data usage. The study highlights a shift towards dynamic, on-policy training methods for better AI alignment. The associated code and data have been made publicly available to support further research in this area.
cs.AI updates on arXiv.org