Path-Coupled Bellman Flows for Distributional Reinforcement Learning
Researchers have introduced Path-Coupled Bellman Flows (PCBF), a novel continuous-time method for Distributional Reinforcement Learning (DRL). Traditional DRL approaches often rely on projections or suffer from boundary mismatches and high-variance bootstrapping in flow-based models. PCBF addresses these limitations by employing source-consistent Bellman-coupled paths, where the current path begins at a base prior and reaches the Bellman target while maintaining an affine relation to successor flows. This technique couples current and successor return flows through shared base noise and utilizes a lambda-parameterized control-variate target to balance bias and variance. Experimental results across analytically tractable Markov Reward Processes, OGBench, and D4RL benchmarks demonstrate that PCBF significantly improves distributional fidelity and training stability. Furthermore, the method achieves competitive performance in offline reinforcement learning tasks. This advancement offers a more robust framework for modeling full return distributions, potentially enhancing the reliability and efficiency of complex AI decision-making systems in various applications.
Wire timeline
Path-Coupled Bellman Flows for Distributional Reinforcement Learning
Researchers have introduced Path-Coupled Bellman Flows (PCBF), a novel continuous-time method for Distributional Reinforcement Learning (DRL). Traditional DRL approaches often rely on projections or suffer from boundary mismatches and high-variance bootstrapping in flow-based models. PCBF addresses these limitations by employing source-consistent Bellman-coupled paths, where the current path begins at a base prior and reaches the Bellman target while maintaining an affine relation to successor flows. This technique couples current and successor return flows through shared base noise and utilizes a lambda-parameterized control-variate target to balance bias and variance. Experimental results across analytically tractable Markov Reward Processes, OGBench, and D4RL benchmarks demonstrate that PCBF significantly improves distributional fidelity and training stability. Furthermore, the method achieves competitive performance in offline reinforcement learning tasks. This advancement offers a more robust framework for modeling full return distributions, potentially enhancing the reliability and efficiency of complex AI decision-making systems in various applications.
cs.AI updates on arXiv.org