Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
Researchers have introduced a new method called Gated Cropped Attention-Delta steering (GCAD) to enhance the reliability of activation steering in large language models. While activation steering typically controls model behavior by modifying internal representations, standard approaches often fail in stateful dialogues due to KV-cache contamination. This issue causes steered token states to be stored and reused, leading to cumulative coherence degradation over time. The proposed GCAD technique addresses this by extracting steering signals from system-prompt contributions to self-attention and applying them with token-level gating. Experimental results demonstrate that GCAD significantly improves long-horizon coherence while preserving trait control. In multi-turn benchmarks, the method reduced average coherence drift from -18.6 to -1.9 and increased turn-10 trait expression from 78.0 to 93.1. These findings suggest that aligning interventions with prompt-mediated pathways used by models for behavioral control yields more robust performance. The study highlights a critical advancement in managing AI behavior during extended interactions, offering a solution to persistent issues in maintaining consistency and persona adherence in complex dialogue systems.
Wire timeline
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
Researchers have introduced a new method called Gated Cropped Attention-Delta steering (GCAD) to enhance the reliability of activation steering in large language models. While activation steering typically controls model behavior by modifying internal representations, standard approaches often fail in stateful dialogues due to KV-cache contamination. This issue causes steered token states to be stored and reused, leading to cumulative coherence degradation over time. The proposed GCAD technique addresses this by extracting steering signals from system-prompt contributions to self-attention and applying them with token-level gating. Experimental results demonstrate that GCAD significantly improves long-horizon coherence while preserving trait control. In multi-turn benchmarks, the method reduced average coherence drift from -18.6 to -1.9 and increased turn-10 trait expression from 78.0 to 93.1. These findings suggest that aligning interventions with prompt-mediated pathways used by models for behavioral control yields more robust performance. The study highlights a critical advancement in managing AI behavior during extended interactions, offering a solution to persistent issues in maintaining consistency and persona adherence in complex dialogue systems.
cs.AI updates on arXiv.org