PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning
Researchers have introduced PiCA, a novel reinforcement learning mechanism designed to enhance Large Language Model (LLM)-based search agents in knowledge-intensive tasks. Current methods struggle with long-horizon credit assignment due to reward sparsity, isolated credit allocation, and distributional shifts. PiCA addresses these challenges by reformulating search trajectories as cumulative progress processes. It utilizes Potential-Based Reward Shaping to define process rewards based on historical context, identifying pivotal steps—specifically target golden sub-queries and sub-answers—as information peaks that significantly increase the probability of correct final answers. This approach provides dense, trajectory-dependent guidance while maintaining distributional consistency. Extensive experiments across seven benchmarks demonstrate that PiCA outperforms existing baselines, achieving performance improvements of 15.2% for 3B models and 2.2% for 7B models. The results highlight PiCA's robust generalization capabilities across various model sizes. The research team has made the code publicly available to support further development in agentic reinforcement learning.
Wire timeline
PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning
Researchers have introduced PiCA, a novel reinforcement learning mechanism designed to enhance Large Language Model (LLM)-based search agents in knowledge-intensive tasks. Current methods struggle with long-horizon credit assignment due to reward sparsity, isolated credit allocation, and distributional shifts. PiCA addresses these challenges by reformulating search trajectories as cumulative progress processes. It utilizes Potential-Based Reward Shaping to define process rewards based on historical context, identifying pivotal steps—specifically target golden sub-queries and sub-answers—as information peaks that significantly increase the probability of correct final answers. This approach provides dense, trajectory-dependent guidance while maintaining distributional consistency. Extensive experiments across seven benchmarks demonstrate that PiCA outperforms existing baselines, achieving performance improvements of 15.2% for 3B models and 2.2% for 7B models. The results highlight PiCA's robust generalization capabilities across various model sizes. The research team has made the code publicly available to support further development in agentic reinforcement learning.
cs.AI updates on arXiv.org