RuPLaR: Efficient Latent Compression of LLM Reasoning Chains with Rule-Based Priors
Researchers have introduced RuPLaR, a novel framework designed to enhance the efficiency and accuracy of Large Language Model (LLM) reasoning. Addressing the limitations of traditional Chain-of-Thought methods, which suffer from natural language inefficiencies, and existing latent approaches plagued by structural complexities, RuPLaR employs a One-Model One-Step architecture. This method trains an LLM to autonomously generate latent reasoning tokens in a single stage, guided by rule-based prior probability distributions. By eliminating cascaded processes and inter-model dependencies, the framework reduces error propagation and coordination overhead. The training objective integrates cross-entropy for answer consistency, KL divergence for aligning soft tokens with priors, and semantic alignment constraints. Extensive experiments demonstrate that RuPLaR improves accuracy by 11.1% over current latent CoT methods while significantly minimizing token usage. This advancement highlights a significant step forward in optimizing LLM reasoning capabilities, offering a more effective and extensible solution for complex problem-solving tasks in artificial intelligence.
Wire timeline
RuPLaR: Efficient Latent Compression of LLM Reasoning Chains with Rule-Based Priors
Researchers have introduced RuPLaR, a novel framework designed to enhance the efficiency and accuracy of Large Language Model (LLM) reasoning. Addressing the limitations of traditional Chain-of-Thought methods, which suffer from natural language inefficiencies, and existing latent approaches plagued by structural complexities, RuPLaR employs a One-Model One-Step architecture. This method trains an LLM to autonomously generate latent reasoning tokens in a single stage, guided by rule-based prior probability distributions. By eliminating cascaded processes and inter-model dependencies, the framework reduces error propagation and coordination overhead. The training objective integrates cross-entropy for answer consistency, KL divergence for aligning soft tokens with priors, and semantic alignment constraints. Extensive experiments demonstrate that RuPLaR improves accuracy by 11.1% over current latent CoT methods while significantly minimizing token usage. This advancement highlights a significant step forward in optimizing LLM reasoning capabilities, offering a more effective and extensible solution for complex problem-solving tasks in artificial intelligence.
cs.AI updates on arXiv.org