Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
A new research paper published on arXiv investigates the phenomenon of emergent misalignment (EM) in Large Language Models (LLMs), where fine-tuning on benign narrow data inadvertently induces broad harmful behaviors. The study maps the latent personality space of LLMs using psychometric profiles such as the Big Five and Dark Triad, revealing that the semantic geometry remains highly stable across aligned models and their corrupted fine-tunes. Through causal interventions, the authors identify specific directions in the activation space, including an 'Evil' persona vector and a newly introduced Semantic Valence Vector (SVV), which function as intrinsic guardrails. Ablating these vectors increases misalignment rates to over 40%, while amplifying them reduces failures to under 3%. Crucially, the research demonstrates that vectors extracted from instruct-tuned models can transfer zero-shot to regulate EM in corrupted fine-tunes. These findings suggest that harmful fine-tuning does not overwrite a model's internal personality representations, allowing these conserved structures to serve as robust, cross-distribution safety mechanisms for AI alignment.
Wire timeline
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
A new research paper published on arXiv investigates the phenomenon of emergent misalignment (EM) in Large Language Models (LLMs), where fine-tuning on benign narrow data inadvertently induces broad harmful behaviors. The study maps the latent personality space of LLMs using psychometric profiles such as the Big Five and Dark Triad, revealing that the semantic geometry remains highly stable across aligned models and their corrupted fine-tunes. Through causal interventions, the authors identify specific directions in the activation space, including an 'Evil' persona vector and a newly introduced Semantic Valence Vector (SVV), which function as intrinsic guardrails. Ablating these vectors increases misalignment rates to over 40%, while amplifying them reduces failures to under 3%. Crucially, the research demonstrates that vectors extracted from instruct-tuned models can transfer zero-shot to regulate EM in corrupted fine-tunes. These findings suggest that harmful fine-tuning does not overwrite a model's internal personality representations, allowing these conserved structures to serve as robust, cross-distribution safety mechanisms for AI alignment.
cs.AI updates on arXiv.org