Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws
A new research paper titled 'The Geometric Wall' challenges the assumption that activation spaces in large language models are globally linear. The study argues that Sparse Autoencoders (SAEs) face a geometry-dependent performance limit, or 'wall,' rather than a finite-resource ceiling. By analyzing 844 residual-stream Gemma Scope SAE checkpoints across 68 layers of Gemma 2 models (2B and 9B parameters), the authors demonstrate that reconstruction error varies significantly across layers due to manifold curvature and intrinsic dimension. The research introduces a two-stage methodology: fitting per-layer scaling-law surfaces and regressing parameters against geometric summaries. Results indicate that manifold geometry accurately predicts per-layer width exponents, with regression coefficients transferring successfully between different model sizes. Higher curvature and intrinsic dimensions correspond to higher asymptotic error floors, confirming that sparse linear approximations struggle with curved manifolds. This work establishes a transferable geometric law for SAE scaling, offering deeper insights into the structural limitations of interpretability tools in artificial intelligence.
Wire timeline
Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws
A new research paper titled 'The Geometric Wall' challenges the assumption that activation spaces in large language models are globally linear. The study argues that Sparse Autoencoders (SAEs) face a geometry-dependent performance limit, or 'wall,' rather than a finite-resource ceiling. By analyzing 844 residual-stream Gemma Scope SAE checkpoints across 68 layers of Gemma 2 models (2B and 9B parameters), the authors demonstrate that reconstruction error varies significantly across layers due to manifold curvature and intrinsic dimension. The research introduces a two-stage methodology: fitting per-layer scaling-law surfaces and regressing parameters against geometric summaries. Results indicate that manifold geometry accurately predicts per-layer width exponents, with regression coefficients transferring successfully between different model sizes. Higher curvature and intrinsic dimensions correspond to higher asymptotic error floors, confirming that sparse linear approximations struggle with curved manifolds. This work establishes a transferable geometric law for SAE scaling, offering deeper insights into the structural limitations of interpretability tools in artificial intelligence.
cs.AI updates on arXiv.org