Where Do Reasoning Models Refuse? Analyzing Safety Mechanisms in AI Chain-of-Thought
This research paper investigates the safety mechanisms of reasoning-based large language models, specifically examining where and how these models decide to refuse harmful requests. Unlike standard chat models that must decide on refusal before generating any output, reasoning models produce extended chains of thought (CoT) prior to their final response. The study analyzes four open-source reasoning models to determine if the CoT process causally influences refusal outcomes. Results indicate that fixing a specific reasoning trace significantly reduces variance in refusal decisions. In distilled models, subtle variations in the opening sentence of the CoT can fully determine whether the model refuses or complies, with these patterns transferring across models from the same teacher. Furthermore, the authors extracted linear refusal directions from model activations. Ablating these directions increased harmful compliance, although this method proved less reliable than in non-reasoning models and caused noticeable degradation in general capabilities. The findings highlight the complex interplay between reasoning processes and safety alignment in advanced AI systems.
Wire timeline
Where Do Reasoning Models Refuse? Analyzing Safety Mechanisms in AI Chain-of-Thought
This research paper investigates the safety mechanisms of reasoning-based large language models, specifically examining where and how these models decide to refuse harmful requests. Unlike standard chat models that must decide on refusal before generating any output, reasoning models produce extended chains of thought (CoT) prior to their final response. The study analyzes four open-source reasoning models to determine if the CoT process causally influences refusal outcomes. Results indicate that fixing a specific reasoning trace significantly reduces variance in refusal decisions. In distilled models, subtle variations in the opening sentence of the CoT can fully determine whether the model refuses or complies, with these patterns transferring across models from the same teacher. Furthermore, the authors extracted linear refusal directions from model activations. Ablating these directions increased harmful compliance, although this method proved less reliable than in non-reasoning models and caused noticeable degradation in general capabilities. The findings highlight the complex interplay between reasoning processes and safety alignment in advanced AI systems.
cs.AI updates on arXiv.org