New Semantic Information Theory Proposes TOKEN as Atomic Unit for LLMs
A new academic treatise titled 'Forget BIT, It is All about TOKEN' introduces a Semantic Information Theory for Large Language Models (LLMs), aiming to replace the heuristic, experiment-driven paradigm with a first-principles theoretical framework. The paper argues for a paradigm shift from the classical BIT, which lacks semantic content, to the TOKEN as the macroscopic atomic carrier of meaning and reasoning. By integrating statistical physics, signal processing, and classical information theory, the authors recast Transformers as energy-based models and interpret semantic embeddings on a semantic manifold. The study models LLMs as stateful channels with feedback, utilizing Massey's directed information to measure autoregressive generation causally. This approach yields a directed rate-distortion function for pre-training and a directed rate-reward function for reinforcement learning post-training. Furthermore, it aligns next-token prediction with Granger causal inference and evaluates LLM reasoning limits against Pearl's Ladder of Causation. The work posits that while the BIT defined the Information Epoch, the TOKEN will define the emerging AI Epoch, offering a rigorous mathematical foundation for understanding semantic information flow in artificial intelligence systems.
Wire timeline
New Semantic Information Theory Proposes TOKEN as Atomic Unit for LLMs
A new academic treatise titled 'Forget BIT, It is All about TOKEN' introduces a Semantic Information Theory for Large Language Models (LLMs), aiming to replace the heuristic, experiment-driven paradigm with a first-principles theoretical framework. The paper argues for a paradigm shift from the classical BIT, which lacks semantic content, to the TOKEN as the macroscopic atomic carrier of meaning and reasoning. By integrating statistical physics, signal processing, and classical information theory, the authors recast Transformers as energy-based models and interpret semantic embeddings on a semantic manifold. The study models LLMs as stateful channels with feedback, utilizing Massey's directed information to measure autoregressive generation causally. This approach yields a directed rate-distortion function for pre-training and a directed rate-reward function for reinforcement learning post-training. Furthermore, it aligns next-token prediction with Granger causal inference and evaluates LLM reasoning limits against Pearl's Ladder of Causation. The work posits that while the BIT defined the Information Epoch, the TOKEN will define the emerging AI Epoch, offering a rigorous mathematical foundation for understanding semantic information flow in artificial intelligence systems.
cs.AI updates on arXiv.org