Learning Strategic Value and Cooperation in Multi-Player Stochastic Games through Side Payments
This academic paper investigates general-sum, multi-player stochastic games where agents utilize transferable utility and side payments to facilitate individually rational cooperation. Building upon the Harsanyi-Shapley (HS) value concept from normal-form games, the authors introduce two new value notions for stochastic settings: HS-S, which aggregates dynamic coalition threat powers, and Coco-S, defined as fixed points of a statewise HS Bellman operator. The study extends HS-style axioms to stochastic environments, proving that HS-S is the unique mapping satisfying these conditions. While HS-S and Coco-S coincide in two-player games, the authors demonstrate via a three-player counterexample that they diverge when more than two players are involved. The research establishes the existence and uniqueness of Coco-S fixed points for specific game configurations using topological degree theory and introduces a Markov Consistency axiom to characterize Coco-S. Additionally, the paper provides sampling-based estimators with finite-sample guarantees and empirically evaluates the induced values, policies, and side payments using multi-player grid-game benchmarks, contributing significantly to computer science and game theory.
Wire timeline
Learning Strategic Value and Cooperation in Multi-Player Stochastic Games through Side Payments
This academic paper investigates general-sum, multi-player stochastic games where agents utilize transferable utility and side payments to facilitate individually rational cooperation. Building upon the Harsanyi-Shapley (HS) value concept from normal-form games, the authors introduce two new value notions for stochastic settings: HS-S, which aggregates dynamic coalition threat powers, and Coco-S, defined as fixed points of a statewise HS Bellman operator. The study extends HS-style axioms to stochastic environments, proving that HS-S is the unique mapping satisfying these conditions. While HS-S and Coco-S coincide in two-player games, the authors demonstrate via a three-player counterexample that they diverge when more than two players are involved. The research establishes the existence and uniqueness of Coco-S fixed points for specific game configurations using topological degree theory and introduces a Markov Consistency axiom to characterize Coco-S. Additionally, the paper provides sampling-based estimators with finite-sample guarantees and empirically evaluates the induced values, policies, and side payments using multi-player grid-game benchmarks, contributing significantly to computer science and game theory.
cs.AI updates on arXiv.org