Learning Approximate Nash Equilibria in Cooperative Multi-Agent Reinforcement Learning via Mean-Field Subsampling
Researchers Emile Anand and Ishani Karmarkar have published a study on arXiv addressing decision-making in large-scale platforms and networked control systems with strict observability constraints. The paper investigates a cooperative Markov game featuring a centralized global agent and numerous homogeneous local agents within a communication-constrained environment. In this setup, the global agent observes only a subset of local agent states at each time step. The authors propose a novel alternating learning framework called ALTERNATING-MARL. This method involves the global agent performing subsampled mean-field Q-learning against a fixed local policy, while local agents update their strategies by optimizing within an induced Markov Decision Process. The study theoretically proves that these approximate best-response dynamics converge to an approximate Nash Equilibrium with an error bound related to the sample size. Furthermore, the approach effectively separates sample complexities between joint state and action spaces. The validity of the proposed framework is demonstrated through numerical simulations focused on multi-robot control applications, offering potential advancements for scalable multi-agent systems.
Wire timeline
Learning Approximate Nash Equilibria in Cooperative Multi-Agent Reinforcement Learning via Mean-Field Subsampling
Researchers Emile Anand and Ishani Karmarkar have published a study on arXiv addressing decision-making in large-scale platforms and networked control systems with strict observability constraints. The paper investigates a cooperative Markov game featuring a centralized global agent and numerous homogeneous local agents within a communication-constrained environment. In this setup, the global agent observes only a subset of local agent states at each time step. The authors propose a novel alternating learning framework called ALTERNATING-MARL. This method involves the global agent performing subsampled mean-field Q-learning against a fixed local policy, while local agents update their strategies by optimizing within an induced Markov Decision Process. The study theoretically proves that these approximate best-response dynamics converge to an approximate Nash Equilibrium with an error bound related to the sample size. Furthermore, the approach effectively separates sample complexities between joint state and action spaces. The validity of the proposed framework is demonstrated through numerical simulations focused on multi-robot control applications, offering potential advancements for scalable multi-agent systems.
cs.AI updates on arXiv.org