Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
A new research paper published on arXiv investigates the fundamental structural vulnerabilities that allow aligned large language models (LLMs) to remain susceptible to jailbreak attacks. The study introduces the concept of Refusal-Escape Directions (RED), defined as local perturbation directions around harmful inputs that shift model behavior from refusal to answering while maintaining harmful semantics. The authors theoretically prove that RED can be decomposed into contributions from specific operator-level sources, including normalization, residual-wiring, and terminal sources. The research highlights a conditional safety-utility trade-off, noting that eliminating these vulnerabilities requires shared expressive modules like self-attention and MLPs to remove constrained source contributions without compromising benign response mechanisms. Empirical experiments across multiple models demonstrate that added token dimensions can expose RED, and successful jailbreaks align significantly with terminal-source contributions. This work provides a mechanistic understanding of why discrete prompt constructions succeed by framing them as continuous input transformations, offering critical insights for improving LLM safety architectures.
Wire timeline
Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
A new research paper published on arXiv investigates the fundamental structural vulnerabilities that allow aligned large language models (LLMs) to remain susceptible to jailbreak attacks. The study introduces the concept of Refusal-Escape Directions (RED), defined as local perturbation directions around harmful inputs that shift model behavior from refusal to answering while maintaining harmful semantics. The authors theoretically prove that RED can be decomposed into contributions from specific operator-level sources, including normalization, residual-wiring, and terminal sources. The research highlights a conditional safety-utility trade-off, noting that eliminating these vulnerabilities requires shared expressive modules like self-attention and MLPs to remove constrained source contributions without compromising benign response mechanisms. Empirical experiments across multiple models demonstrate that added token dimensions can expose RED, and successful jailbreaks align significantly with terminal-source contributions. This work provides a mechanistic understanding of why discrete prompt constructions succeed by framing them as continuous input transformations, offering critical insights for improving LLM safety architectures.
cs.AI updates on arXiv.org