Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering
A new research paper published on arXiv highlights significant safety vulnerabilities in Large Reasoning Models (LRMs) that expose chain-of-thought reasoning. The study reveals that while final answers may appear safe, the intermediate reasoning traces often contain harmful or policy-violating content, creating a critical safety blind spot. Researchers evaluated 15 open-weight and API-based LRMs using over 41,000 prompts, identifying two primary failure modes: 'leak cases,' where unsafe reasoning precedes a safe answer, and 'escape cases,' where benign reasoning leads to an unsafe response. Risks were concentrated in areas such as misinformation, legal compliance, and physical harm. To address this, the authors propose 'adaptive multi-principle steering,' a white-box mitigation technique that adjusts model activations during inference. Tested on models like DeepSeek-R1-Qwen-7B, this method reduced unsafe content by 40.8% while maintaining 97.7% accuracy on standard benchmarks. The findings suggest that LRM safety evaluations must encompass the entire reasoning trajectory, not just the final output, to ensure robust alignment with safety principles.
Wire timeline
Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering
A new research paper published on arXiv highlights significant safety vulnerabilities in Large Reasoning Models (LRMs) that expose chain-of-thought reasoning. The study reveals that while final answers may appear safe, the intermediate reasoning traces often contain harmful or policy-violating content, creating a critical safety blind spot. Researchers evaluated 15 open-weight and API-based LRMs using over 41,000 prompts, identifying two primary failure modes: 'leak cases,' where unsafe reasoning precedes a safe answer, and 'escape cases,' where benign reasoning leads to an unsafe response. Risks were concentrated in areas such as misinformation, legal compliance, and physical harm. To address this, the authors propose 'adaptive multi-principle steering,' a white-box mitigation technique that adjusts model activations during inference. Tested on models like DeepSeek-R1-Qwen-7B, this method reduced unsafe content by 40.8% while maintaining 97.7% accuracy on standard benchmarks. The findings suggest that LRM safety evaluations must encompass the entire reasoning trajectory, not just the final output, to ensure robust alignment with safety principles.
cs.AI updates on arXiv.org