Cross-Family Universality of Behavioral Axes via Anchor-Projected Representations
Researchers Su-Hyeon Kim and Yo-Us Han introduce a novel anchor-projection framework to address the challenge of comparing behavioral directions across large language models (LLMs) with differing architectures. By mapping hidden representations into a shared anchor coordinate space (ACS), the study enables the transfer of behavioral axes without fine-tuning. Evaluating five instruction-tuned model families, including Llama, Qwen, Mistral, and Phi, the authors find that behavioral directions align tightly within this cluster. The framework achieves high performance in downstream tasks, with held-out targets reaching 0.83 ten-way detection accuracy and 0.95 mean binary AUROC. Additionally, canonical steering effectively induces refusal-rate shifts. Sensitivity analyses indicate that robust transferable directions can be approximated using just two source models and small anchor pools. This work provides significant insights into cross-family interpretability, demonstrating that representation-level transfer remains robust despite variations in tokenizers and training procedures, offering a new perspective on universal behavioral structures in AI.
Wire timeline
Cross-Family Universality of Behavioral Axes via Anchor-Projected Representations
Researchers Su-Hyeon Kim and Yo-Us Han introduce a novel anchor-projection framework to address the challenge of comparing behavioral directions across large language models (LLMs) with differing architectures. By mapping hidden representations into a shared anchor coordinate space (ACS), the study enables the transfer of behavioral axes without fine-tuning. Evaluating five instruction-tuned model families, including Llama, Qwen, Mistral, and Phi, the authors find that behavioral directions align tightly within this cluster. The framework achieves high performance in downstream tasks, with held-out targets reaching 0.83 ten-way detection accuracy and 0.95 mean binary AUROC. Additionally, canonical steering effectively induces refusal-rate shifts. Sensitivity analyses indicate that robust transferable directions can be approximated using just two source models and small anchor pools. This work provides significant insights into cross-family interpretability, demonstrating that representation-level transfer remains robust despite variations in tokenizers and training procedures, offering a new perspective on universal behavioral structures in AI.
cs.AI updates on arXiv.org