Internalizing Safety Understanding in Large Reasoning Models via Verification
Researchers from the University of Science and Technology of China have introduced Safety Internal (SInternal), a novel framework designed to enhance the safety of Large Reasoning Models (LRMs). While Chain-of-Thought reasoning empowers these models, it often leads to riskier outputs. Current alignment methods primarily rely on external compliance, focusing on detecting malicious prompts rather than evaluating output safety, leaving models vulnerable to adversarial jailbreaks. The SInternal framework addresses this by training LRMs exclusively on safety verification tasks, enabling them to critique their own generated answers using expert reasoning trajectories. Empirical analysis demonstrates that this approach internalizes safety specifications, fostering intrinsic safety understanding rather than mere behavioral mimicry. The method significantly improves robustness against out-of-domain jailbreaks and serves as a superior initialization for reinforcement learning compared to standard supervised fine-tuning. This research suggests that internalizing safety understanding creates a more robust foundation for AI alignment. The associated code has been made publicly available to support further development and verification in the field of artificial intelligence safety.
Wire timeline
Internalizing Safety Understanding in Large Reasoning Models via Verification
Researchers from the University of Science and Technology of China have introduced Safety Internal (SInternal), a novel framework designed to enhance the safety of Large Reasoning Models (LRMs). While Chain-of-Thought reasoning empowers these models, it often leads to riskier outputs. Current alignment methods primarily rely on external compliance, focusing on detecting malicious prompts rather than evaluating output safety, leaving models vulnerable to adversarial jailbreaks. The SInternal framework addresses this by training LRMs exclusively on safety verification tasks, enabling them to critique their own generated answers using expert reasoning trajectories. Empirical analysis demonstrates that this approach internalizes safety specifications, fostering intrinsic safety understanding rather than mere behavioral mimicry. The method significantly improves robustness against out-of-domain jailbreaks and serves as a superior initialization for reinforcement learning compared to standard supervised fine-tuning. This research suggests that internalizing safety understanding creates a more robust foundation for AI alignment. The associated code has been made publicly available to support further development and verification in the field of artificial intelligence safety.
cs.AI updates on arXiv.org