Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
Researchers Nilesh Sarkar and Dawar Jyoti Deka introduce a new metric, causal dimensionality kappa, to characterize the relationship between Sparse Autoencoder (SAE) width and the causal influence of transformer representations on model output. Published on arXiv, the study defines kappa as the effective rank of the expected Jacobian outer product at specific layers. Through experiments on Gemma-2 models, the authors identify a 'representational-causal wedge,' where representational capacity grows significantly faster than causal capacity as SAE width increases. Key findings indicate that kappa is invariant to model scaling, with Gemma-2-9B and Gemma-2-2B showing identical causal metrics despite parameter differences. Additionally, kappa remains constant across network depths, although attribution thresholds decrease substantially from early to late layers. The research establishes kappa as a measurable, intrinsic property of transformer layers, offering insights into how interpretability features scale and structure within large language models. This work provides a systematic method for estimating causal influence via SAE width sweeps and attribution patching, contributing to the understanding of mechanistic interpretability in artificial intelligence.
Wire timeline
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
Researchers Nilesh Sarkar and Dawar Jyoti Deka introduce a new metric, causal dimensionality kappa, to characterize the relationship between Sparse Autoencoder (SAE) width and the causal influence of transformer representations on model output. Published on arXiv, the study defines kappa as the effective rank of the expected Jacobian outer product at specific layers. Through experiments on Gemma-2 models, the authors identify a 'representational-causal wedge,' where representational capacity grows significantly faster than causal capacity as SAE width increases. Key findings indicate that kappa is invariant to model scaling, with Gemma-2-9B and Gemma-2-2B showing identical causal metrics despite parameter differences. Additionally, kappa remains constant across network depths, although attribution thresholds decrease substantially from early to late layers. The research establishes kappa as a measurable, intrinsic property of transformer layers, offering insights into how interpretability features scale and structure within large language models. This work provides a systematic method for estimating causal influence via SAE width sweeps and attribution patching, contributing to the understanding of mechanistic interpretability in artificial intelligence.
cs.AI updates on arXiv.org