English

Model and Reinforcement Learning for Markov Games with Risk Preferences

Computer Science and Game Theory 2019-11-22 v2 Multiagent Systems Optimization and Control

Abstract

We motivate and propose a new model for non-cooperative Markov game which considers the interactions of risk-aware players. This model characterizes the time-consistent dynamic "risk" from both stochastic state transitions (inherent to the game) and randomized mixed strategies (due to all other players). An appropriate risk-aware equilibrium concept is proposed and the existence of such equilibria is demonstrated in stationary strategies by an application of Kakutani's fixed point theorem. We further propose a simulation-based Q-learning type algorithm for risk-aware equilibrium computation. This algorithm works with a special form of minimax risk measures which can naturally be written as saddle-point stochastic optimization problems, and covers many widely investigated risk measures. Finally, the almost sure convergence of this simulation-based algorithm to an equilibrium is demonstrated under some mild conditions. Our numerical experiments on a two player queuing game validate the properties of our model and algorithm, and demonstrate their worth and applicability in real life competitive decision-making.

Keywords

Cite

@article{arxiv.1901.04882,
  title  = {Model and Reinforcement Learning for Markov Games with Risk Preferences},
  author = {Wenjie Huang and Pham Viet Hai and William B. Haskell},
  journal= {arXiv preprint arXiv:1901.04882},
  year   = {2019}
}

Comments

38 pages, 6 tables, 5 figures

R2 v1 2026-06-23T07:12:28.541Z