Related papers: Actor-Dual-Critic Dynamics for Zero-sum and Identi…
In many multi-player interactions, players incur strictly positive costs each time they execute actions e.g. 'menu costs' or transaction costs in financial systems. Since acting at each available opportunity would accumulate prohibitively…
When deployed in the world, a learning agent such as a recommender system or a chatbot often repeatedly interacts with another learning agent (such as a user) over time. In many such two-agent systems, each agent learns separately and the…
In common-interest stochastic games all players receive an identical payoff. Players participating in such games must learn to coordinate with each other in order to receive the highest-possible value. A number of reinforcement learning…
The classical, complete-information two-player games assume that the problem data (in particular the payoff matrix) is known exactly by both players. In a now famous result, Nash has shown that any such game has an equilibrium in mixed…
This paper investigates the convergence of learning dynamics in Stackelberg games. In the class of games we consider, there is a hierarchical game being played between a leader and a follower with continuous action spaces. We establish a…
We consider a zero-sum stochastic game for continuous-time Markov chain with countable state space and unbounded transition and pay-off rates. The additional feature of the game is that the controllers together with taking actions are also…
We study the problem of off-policy critic evaluation in several variants of value-based off-policy actor-critic algorithms. Off-policy actor-critic algorithms require an off-policy critic evaluation step, to estimate the value of the new…
Cooperation in heterogeneous groups, where individuals differ in resources, productivity, and behavioural responsiveness, underpins collective action across many social and biological systems. Introspection dynamics, in which each player…
We develop here the Stochastic Perron Method in the framework of two-player zero-sum differential games. We consider the formulation of the game where both players play, symmetrically, feed-back strategies (as in [CR09] or [PZ12]) as…
We introduce a stochastic learning process called the dampened gradient approximation process. While learning models have almost exclusively focused on finite games, in this paper we design a learning process for games with continuous…
Zero sum games with risk-sensitive cost criterion are considered with underlying dynamics being given by controlled stochastic differential equations. Under the assumption of geometric stability on the dynamics , we completely characterize…
We propose a learning dynamics to model how strategic agents repeatedly play a continuous game while relying on an information platform to learn an unknown payoff-relevant parameter. In each time step, the platform updates a belief estimate…
We study a dynamic sender-receiver game in which the sender observes a state evolving according to a Markov chain but does not observe the receiver's action. Despite the absence of feedback, dynamic interaction partially restores…
This paper is concerned with a two-person zero-sum indefinite stochastic linear-quadratic Stackelberg differential game with asymmetric informational uncertainties, where both the leader and follower face different and unknown disturbances.…
An important challenge in non-cooperative game theory is coordinating on a single (approximate) equilibrium from many possibilities - a challenge that becomes even more complex when players hold private information. Recommender mechanisms…
While multi-agent reinforcement learning (MARL) has produced numerous algorithms that converge to Nash or related equilibria, such equilibria are often non-unique and can exhibit widely varying efficiency. This raises a fundamental…
We study a class of stochastic dynamic games that exhibit strategic complementarities between players; formally, in the games we consider, the payoff of a player has increasing differences between her own state and the empirical…
We propose the first model-free algorithm that achieves low regret performance for decentralized learning in two-player zero-sum tabular stochastic games with infinite-horizon average-reward objective. In decentralized learning, the…
This paper studies a class of strongly monotone games involving non-cooperative agents that optimize their own time-varying cost functions. We assume that the agents can observe other agents' historical actions and choose actions that best…
We add the assumption that players know their opponents' payoff functions and rationality to a model of non-equilibrium learning in signaling games. Agents are born into player roles and play against random opponents every period.…