中文
相关论文

相关论文: Recursive Regret Matching: A General Method for So…

200 篇论文

Regret minimization methods are a powerful tool for learning approximate Nash equilibrium (NE) in two-player zero-sum imperfect information extensive-form games (IIEGs). We consider the problem in the interactive bandit-feedback setting…

机器学习 · 计算机科学 2023-08-21 Linjian Meng , Yang Gao

In zero-sum games, the optimal strategy is well-defined by the Nash equilibrium. However, it is overly conservative when playing against suboptimal opponents and it can not exploit their weaknesses. Limited look-ahead game solving in…

计算机科学与博弈论 · 计算机科学 2024-04-04 David Milec , Ondřej Kubíček , Viliam Lisý

Synthesis of finite-state controllers from high-level specifications in multi-agent systems can be reduced to solving multi-player concurrent games over finite graphs. The complexity of solving such games with qualitative objectives for…

计算机科学与博弈论 · 计算机科学 2018-09-28 Shaull Almagor , Rajeev Alur , Suguman Bansal

Several notions of game enjoy a Nash-like notion of equilibrium without guarantee of existence. There are different ways of weakening a definition of Nash-like equilibrium in order to guarantee the existence of a weakened equilibrium.…

计算机科学与博弈论 · 计算机科学 2007-12-11 Stéphane Le Roux

Extensive-form games with imperfect recall are an important game-theoretic model that allows a compact representation of strategies in dynamic strategic interactions. Practical use of imperfect recall games is limited due to negative…

计算机科学与博弈论 · 计算机科学 2017-05-25 Branislav Bosansky , Jiri Cermak , Karel Horak , Michal Pechoucek

In this paper, we study the problem of learning the set of pure strategy Nash equilibria and the exact structure of a continuous-action graphical game with quadratic payoffs by observing a small set of perturbed equilibria. A…

计算机科学与博弈论 · 计算机科学 2019-11-12 Adarsh Barik , Jean Honorio

One key in real-life Nash equilibrium applications is to calibrate players' cost functions. To leverage the approximation ability of neural networks, we proposed a general framework for optimizing and learning Nash equilibrium using neural…

计算机科学与博弈论 · 计算机科学 2024-09-04 Di Zhang , Wei Gu , Qing Jin

The existence of simple, uncoupled no-regret dynamics that converge to correlated equilibria in normal-form games is a celebrated result in the theory of multi-agent systems. Specifically, it has been known for more than 20 years that when…

计算机科学与博弈论 · 计算机科学 2022-09-05 Andrea Celli , Alberto Marchesi , Gabriele Farina , Nicola Gatti

We consider online learning in multi-player smooth monotone games. Existing algorithms have limitations such as (1) being only applicable to strongly monotone games; (2) lacking the no-regret guarantee; (3) having only asymptotic or slow…

机器学习 · 计算机科学 2023-09-06 Yang Cai , Weiqiang Zheng

In many settings where multiple agents interact, the optimal choices for each agent depend heavily on the choices of the others. These coupled interactions are well-described by a general-sum differential game, in which players have…

机器人学 · 计算机科学 2020-05-07 Lasse Peters , David Fridovich-Keil , Claire J. Tomlin , Zachary N. Sunberg

We show for the first time, to our knowledge, that it is possible to reconcile in online learning in zero-sum games two seemingly contradictory objectives: vanishing time-average regret and non-vanishing step sizes. This phenomenon, that we…

计算机科学与博弈论 · 计算机科学 2019-05-14 James P. Bailey , Georgios Piliouras

We study a setting in which two players play a (possibly approximate) Nash equilibrium of a bimatrix game, while a learner observes only their actions and has no knowledge of the equilibrium or the underlying game. A natural question is…

计算机科学与博弈论 · 计算机科学 2026-05-27 Annalisa Barbara , Riccardo Poiani , Martino Bernasconi , Andrea Celli

Infinitely repeated games can support cooperative outcomes that are not equilibria in the one-shot game. The idea is to make sure that any gains from deviating will be offset by retaliation in future rounds. However, this model of…

计算机科学与博弈论 · 计算机科学 2024-06-04 Ratip Emin Berker , Vincent Conitzer

Continuous games are multiplayer games in which strategy sets are compact and utility functions are continuous. These games typically have a highly complicated structure of Nash equilibria, and numerical methods for the equilibrium…

计算机科学与博弈论 · 计算机科学 2022-07-12 T. Kroupa , T. Votroubek

This paper is an attempt to compute the value and saddle points of zero-sum risk-sensitive average stochastic games. For the average games with finite states and actions, we first introduce the so-called irreducibility coefficient and then…

最优化与控制 · 数学 2025-05-08 Fang Chen , Xianping Guo , Xin Guo , Junyu Zhang

This paper considers data-based solutions of linear-quadratic nonzero-sum differential games. Two cases are considered. First, the deterministic game is solved and Nash equilibrium strategies are obtained by using persistently excited data…

系统与控制 · 电气工程与系统科学 2026-05-15 Victor G. Lopez , Matthias A. Müller

This paper investigates a class of general linear-quadratic mean field games with common noise, where the diffusion terms of the system contain the state variables, control variables, and the average state terms. We solve the problem using…

最优化与控制 · 数学 2025-08-29 Yu Si , Jingtao Shi

In this paper, we present a novel consensus-based zeroth-order algorithm tailored for non-convex multiplayer games. The proposed method leverages a metaheuristic approach using concepts from swarm intelligence to reliably identify global…

动力系统 · 数学 2024-07-30 Enis Chenchene , Hui Huang , Jinniao Qiu

Swap regret is a notion that has proven itself to be central to the study of general-sum normal-form games, with swap-regret minimization leading to convergence to the set of correlated equilibria and guaranteeing non-manipulability against…

计算机科学与博弈论 · 计算机科学 2025-02-28 Eshwar Ram Arunachaleswaran , Natalie Collina , Yishay Mansour , Mehryar Mohri , Jon Schneider , Balasubramanian Sivan

This paper considers no-regret learning for repeated continuous-kernel games with lossy bandit feedback. Since it is difficult to give the explicit model of the utility functions in dynamic environments, the players' action can only be…

机器学习 · 计算机科学 2022-05-17 Wenting Liu , Jinlong Lei , Peng Yi , Yiguang Hong
‹ 上一页 1 8 9 10 下一页 ›