Single Neuron Sufficient to Bypass LLM Safety Alignment
A new research paper published on arXiv reveals a critical vulnerability in the safety alignment of large language models (LLMs). The study, conducted by Hamid Kazemi, Atoosa Chegini, and Maria Safi, demonstrates that safety mechanisms in LLMs rely on two distinct systems: refusal neurons that gate harmful output and concept neurons that encode harmful knowledge. By targeting just one neuron in each system, the researchers successfully bypassed safety protocols across seven different models, ranging from 1.7 billion to 70 billion parameters. They achieved this by suppressing refusal neurons to allow explicit harmful requests or amplifying concept neurons to induce harmful content from innocent prompts. Notably, these exploits required no additional training or prompt engineering. The findings challenge the assumption that safety alignment is robustly distributed across model weights, suggesting instead that individual neurons are causally sufficient to control refusal behavior. This discovery highlights significant security risks in current AI safety frameworks, indicating that existing alignment strategies may be fragile and susceptible to precise, low-effort attacks.
Wire timeline
Single Neuron Sufficient to Bypass LLM Safety Alignment
A new research paper published on arXiv reveals a critical vulnerability in the safety alignment of large language models (LLMs). The study, conducted by Hamid Kazemi, Atoosa Chegini, and Maria Safi, demonstrates that safety mechanisms in LLMs rely on two distinct systems: refusal neurons that gate harmful output and concept neurons that encode harmful knowledge. By targeting just one neuron in each system, the researchers successfully bypassed safety protocols across seven different models, ranging from 1.7 billion to 70 billion parameters. They achieved this by suppressing refusal neurons to allow explicit harmful requests or amplifying concept neurons to induce harmful content from innocent prompts. Notably, these exploits required no additional training or prompt engineering. The findings challenge the assumption that safety alignment is robustly distributed across model weights, suggesting instead that individual neurons are causally sufficient to control refusal behavior. This discovery highlights significant security risks in current AI safety frameworks, indicating that existing alignment strategies may be fragile and susceptible to precise, low-effort attacks.
cs.AI updates on arXiv.org