IMAX Framework Enhances Exploration in RLVR via Prefix-Tuned Priors
Researchers have introduced the Information-Maximizing Augmented eXploration (IMAX) framework to address exploration challenges in Reinforcement Learning with Verifiable Rewards (RLVR) for large language models. Current RLVR methods often suffer from entropy collapse, where models improve single-rollout accuracy but fail to explore diverse successful reasoning trajectories due to reward sparsity and long horizons. Unlike passive techniques like entropy regularization that may degrade generation quality, IMAX trains a pool of soft prefixes to reshape the base model's prior over reasoning paths. Each prefix acts as a trainable control knob, inducing distinct rollout distributions from the same backbone. The framework also incorporates an Information Maximization (InfoMax) reward to encourage diverse, task-relevant behaviors alongside verifiable rewards. Algorithm-agnostic and compatible with existing RLVR pipelines, IMAX demonstrated consistent performance improvements across three backbone scales. Experimental results showed gains of up to 11.60% in Pass@4 and 10.57% in Avg@4 compared to standard RLVR, highlighting its effectiveness in enhancing reasoning capabilities without compromising generation quality.
Wire timeline
IMAX Framework Enhances Exploration in RLVR via Prefix-Tuned Priors
Researchers have introduced the Information-Maximizing Augmented eXploration (IMAX) framework to address exploration challenges in Reinforcement Learning with Verifiable Rewards (RLVR) for large language models. Current RLVR methods often suffer from entropy collapse, where models improve single-rollout accuracy but fail to explore diverse successful reasoning trajectories due to reward sparsity and long horizons. Unlike passive techniques like entropy regularization that may degrade generation quality, IMAX trains a pool of soft prefixes to reshape the base model's prior over reasoning paths. Each prefix acts as a trainable control knob, inducing distinct rollout distributions from the same backbone. The framework also incorporates an Information Maximization (InfoMax) reward to encourage diverse, task-relevant behaviors alongside verifiable rewards. Algorithm-agnostic and compatible with existing RLVR pipelines, IMAX demonstrated consistent performance improvements across three backbone scales. Experimental results showed gains of up to 11.60% in Pass@4 and 10.57% in Avg@4 compared to standard RLVR, highlighting its effectiveness in enhancing reasoning capabilities without compromising generation quality.
cs.AI updates on arXiv.org