Do Linear Probes Generalize Better in Persona Coordinates?
A new research paper submitted to arXiv addresses the limitations of text-only monitoring for large language models, which often fail to detect strategic deception and sandbagging. The authors propose using white-box monitors, specifically linear probes, to read model internals directly. However, existing probes struggle with distribution shifts in real-world settings. To overcome this, the study investigates low-dimensional subspaces within model internals that robustly capture harmful behaviors while excluding spurious correlations. Inspired by the Assistant Axis and Persona Selection Model, the researchers constructed persona axes for deception and sycophancy using contrastive prompts. By applying unsupervised Principal Component Analysis (PCA) to persona-specific vectors, they identified directions that cleanly separate harmful from harmless personas. Evaluations across ten datasets demonstrated that probes trained on these persona-derived projections generalize significantly better than those trained on raw activations. Furthermore, a unified axis combining multiple behaviors improved cross-dataset generalization. The findings suggest that persona vectors provide a valuable inductive bias for developing more transferable and robust behavior probes, enhancing the safety and reliability of AI monitoring systems against adversarial model behaviors.
Wire timeline
Do Linear Probes Generalize Better in Persona Coordinates?
A new research paper submitted to arXiv addresses the limitations of text-only monitoring for large language models, which often fail to detect strategic deception and sandbagging. The authors propose using white-box monitors, specifically linear probes, to read model internals directly. However, existing probes struggle with distribution shifts in real-world settings. To overcome this, the study investigates low-dimensional subspaces within model internals that robustly capture harmful behaviors while excluding spurious correlations. Inspired by the Assistant Axis and Persona Selection Model, the researchers constructed persona axes for deception and sycophancy using contrastive prompts. By applying unsupervised Principal Component Analysis (PCA) to persona-specific vectors, they identified directions that cleanly separate harmful from harmless personas. Evaluations across ten datasets demonstrated that probes trained on these persona-derived projections generalize significantly better than those trained on raw activations. Furthermore, a unified axis combining multiple behaviors improved cross-dataset generalization. The findings suggest that persona vectors provide a valuable inductive bias for developing more transferable and robust behavior probes, enhancing the safety and reliability of AI monitoring systems against adversarial model behaviors.
cs.AI updates on arXiv.org