Harmful Intent Identified as Geometrically Recoverable Feature in LLM Residual Streams
A new research paper published on arXiv demonstrates that harmful intent in large language models (LLMs) is linearly separable within residual-stream activations. The study analyzed 12 models across four architectural families, including Qwen, Llama, and Gemma, ranging from 0.5B to 9B parameters. Using a direction fitted from just 100 labeled examples, the researchers achieved a mean effective AUROC of 0.982, with strong generalization to held-out benchmarks. Notably, this detection capability persists even in abliterated models where refusal mechanisms have been removed, suggesting that harm recognition is distinct from refusal behavior. The findings highlight significant variations in true positive rates at low false positive rates among different supervised strategies. Geometric analysis reveals that the recovered harm direction is protocol-dependent, varying significantly between extraction methods like max-pooling and last-token pooling. This indicates that probing identifies protocol-specific directions rather than a single unique computational feature. The results advocate for low false-positive rate reporting as a standard in safety evaluation, providing deeper insights into how aligned models internally represent and process harmful instructions.
Wire timeline
Harmful Intent Identified as Geometrically Recoverable Feature in LLM Residual Streams
A new research paper published on arXiv demonstrates that harmful intent in large language models (LLMs) is linearly separable within residual-stream activations. The study analyzed 12 models across four architectural families, including Qwen, Llama, and Gemma, ranging from 0.5B to 9B parameters. Using a direction fitted from just 100 labeled examples, the researchers achieved a mean effective AUROC of 0.982, with strong generalization to held-out benchmarks. Notably, this detection capability persists even in abliterated models where refusal mechanisms have been removed, suggesting that harm recognition is distinct from refusal behavior. The findings highlight significant variations in true positive rates at low false positive rates among different supervised strategies. Geometric analysis reveals that the recovered harm direction is protocol-dependent, varying significantly between extraction methods like max-pooling and last-token pooling. This indicates that probing identifies protocol-specific directions rather than a single unique computational feature. The results advocate for low false-positive rate reporting as a standard in safety evaluation, providing deeper insights into how aligned models internally represent and process harmful instructions.
cs.AI updates on arXiv.org