Scaling Limits of Long-Context Transformers
A new academic paper submitted to arXiv investigates the theoretical scaling limits of softmax self-attention mechanisms in long-context Transformers. The study analyzes a fixed query against a random context of independent and identically distributed keys on a sphere, using inverse temperature as a critical scaling parameter. The authors demonstrate that the emergence of selectivity is determined by the local exponent of the distance-to-query distribution near zero, rather than global context features. Specifically, for uniform keys, the critical scale follows a power law relative to the context size and dimension. The research characterizes limiting laws across three distinct regimes: subcritical, where output approximates a local average with Gaussian fluctuations; critical, where nearest keys retain macroscopic mass without collapsing to a single key; and supercritical, where attention concentrates entirely on the closest key. Notably, the subcritical regime with an identity value matrix is shown to approximate a backward heat equation. This work provides rigorous mathematical insights into how attention mechanisms behave as context lengths increase, offering foundational knowledge for optimizing large language models.
Wire timeline
Scaling Limits of Long-Context Transformers
A new academic paper submitted to arXiv investigates the theoretical scaling limits of softmax self-attention mechanisms in long-context Transformers. The study analyzes a fixed query against a random context of independent and identically distributed keys on a sphere, using inverse temperature as a critical scaling parameter. The authors demonstrate that the emergence of selectivity is determined by the local exponent of the distance-to-query distribution near zero, rather than global context features. Specifically, for uniform keys, the critical scale follows a power law relative to the context size and dimension. The research characterizes limiting laws across three distinct regimes: subcritical, where output approximates a local average with Gaussian fluctuations; critical, where nearest keys retain macroscopic mass without collapsing to a single key; and supercritical, where attention concentrates entirely on the closest key. Notably, the subcritical regime with an identity value matrix is shown to approximate a backward heat equation. This work provides rigorous mathematical insights into how attention mechanisms behave as context lengths increase, offering foundational knowledge for optimizing large language models.
cs.AI updates on arXiv.org