Geometric Analysis Reveals Three Phases in LLM Next-Token Prediction
A new research paper published on arXiv investigates the geometry of predictive information within large language models (LLMs). The authors, Gianfranco Lombardo, Giuseppe Trimigno, and Stefano Cagnoni, utilize representation lenses as geometric diagnostic tools to track how predictive information evolves across model layers. By analyzing eight models from the Qwen2.5 and OLMo2 families, ranging from 1B to 32B parameters, the study identifies three distinct geometric phases in next-token prediction. The first phase, Seeding Multiplexing, involves feed-forward and attention layers seeding candidate sets. The second, Hoisting Overriding, concentrates the candidate distribution without expanding rank. The final phase, Focal Convergence, aligns the winning token with the unembedding direction. The findings suggest that updates remain approximately orthogonal to the residual stream, with deeper models primarily using additional capacity for candidate disambiguation. This structural analysis provides new insights into the internal mechanics of transformer architectures, highlighting a consistent Pareto frontier between visibility and energy retention across different scales.
Wire timeline
Geometric Analysis Reveals Three Phases in LLM Next-Token Prediction
A new research paper published on arXiv investigates the geometry of predictive information within large language models (LLMs). The authors, Gianfranco Lombardo, Giuseppe Trimigno, and Stefano Cagnoni, utilize representation lenses as geometric diagnostic tools to track how predictive information evolves across model layers. By analyzing eight models from the Qwen2.5 and OLMo2 families, ranging from 1B to 32B parameters, the study identifies three distinct geometric phases in next-token prediction. The first phase, Seeding Multiplexing, involves feed-forward and attention layers seeding candidate sets. The second, Hoisting Overriding, concentrates the candidate distribution without expanding rank. The final phase, Focal Convergence, aligns the winning token with the unembedding direction. The findings suggest that updates remain approximately orthogonal to the residual stream, with deeper models primarily using additional capacity for candidate disambiguation. This structural analysis provides new insights into the internal mechanics of transformer architectures, highlighting a consistent Pareto frontier between visibility and energy retention across different scales.
cs.AI updates on arXiv.org