Policy Gradient Methods for Non-Markovian Reinforcement Learning
Researchers have introduced a novel approach to reinforcement learning in non-Markovian decision processes (NMDPs), where observations and rewards rely on entire interaction histories. The study proposes the Agent State-Markov Policy Gradient (ASMPG) algorithm, which jointly optimizes agent state dynamics and control policies to maximize expected cumulative reward. Unlike previous methods that treat state dynamics as fixed or learn them via predictive objectives, this reward-centric formulation utilizes an internal state recursively updated to summarize past interactions. The authors establish a new policy gradient theorem for Agent State-Markov (ASM) policies, extending classical results to episodic and infinite-horizon discounted NMDPs. The ASMPG algorithm leverages the recursive structure of these dynamics for efficient optimization. Theoretical analysis provides finite-time and almost sure convergence guarantees. Empirical evaluations demonstrate that ASMPG outperforms baseline methods relying on predictive objectives across various non-Markovian tasks. This work represents a significant advancement in handling complex dependencies in reinforcement learning environments, offering robust theoretical foundations and improved practical performance for agents operating in history-dependent settings.
Wire timeline
Policy Gradient Methods for Non-Markovian Reinforcement Learning
Researchers have introduced a novel approach to reinforcement learning in non-Markovian decision processes (NMDPs), where observations and rewards rely on entire interaction histories. The study proposes the Agent State-Markov Policy Gradient (ASMPG) algorithm, which jointly optimizes agent state dynamics and control policies to maximize expected cumulative reward. Unlike previous methods that treat state dynamics as fixed or learn them via predictive objectives, this reward-centric formulation utilizes an internal state recursively updated to summarize past interactions. The authors establish a new policy gradient theorem for Agent State-Markov (ASM) policies, extending classical results to episodic and infinite-horizon discounted NMDPs. The ASMPG algorithm leverages the recursive structure of these dynamics for efficient optimization. Theoretical analysis provides finite-time and almost sure convergence guarantees. Empirical evaluations demonstrate that ASMPG outperforms baseline methods relying on predictive objectives across various non-Markovian tasks. This work represents a significant advancement in handling complex dependencies in reinforcement learning environments, offering robust theoretical foundations and improved practical performance for agents operating in history-dependent settings.
cs.AI updates on arXiv.org