中文
相关论文

相关论文: Combining Deep Reinforcement Learning and Search f…

200 篇论文

In general, two-agent decision-making problems can be modeled as a two-player game, and a typical solution is to find a Nash equilibrium in such game. Counterfactual regret minimization (CFR) is a well-known method to find a Nash…

计算机科学与博弈论 · 计算机科学 2020-12-07 Huale Li , Xuan Wang , Shuhan Qi , Jiajia Zhang , Yang Liu , Yulin Wu , Fengwei Jia

This paper considers convex games involving multiple agents that aim to minimize their own cost functions using locally available information. A common assumption in the study of such games is that the agents are symmetric, meaning that…

最优化与控制 · 数学 2025-09-25 Zifan Wang , Xinlei Yi , Yi Shen , Michael M. Zavlanos , Karl H. Johansson

We address payoff-based decentralized learning in infinite-horizon zero-sum Markov games. In this setting, each player makes decisions based solely on received rewards, without observing the opponent's strategy or actions nor sharing…

计算机科学与博弈论 · 计算机科学 2025-02-11 Reda Ouhamma , Maryam Kamgarpour

This paper introduces risk-revising players to a class of games with incomplete information. These players enter the game with ex ante risk preferences represented by coherent risk measures and develop time-consistent interim revisions of…

最优化与控制 · 数学 2026-03-23 Shutian Liu

We consider finite-horizon and infinite-horizon versions of a dynamic game with $N$ selfish players who observe their types privately and take actions that are publicly observed. Players' types evolve as conditionally independent Markov…

最优化与控制 · 数学 2018-03-20 Deepanshu Vasal , Abhinav Sinha , Achilleas Anastasopoulos

Unlike Poker where the action space $\mathcal{A}$ is discrete, differential games in the physical world often have continuous action spaces not amenable to discrete abstraction, rendering no-regret algorithms with…

计算机科学与博弈论 · 计算机科学 2025-02-17 Mukesh Ghimire , Zhe Xu , Yi Ren

We consider learning to play multiplayer imperfect-information games with simultaneous moves and large state-action spaces. Previous attempts to tackle such challenging games have largely focused on model-free learning methods, often…

人工智能 · 计算机科学 2020-12-23 Rinu Boney , Alexander Ilin , Juho Kannala , Jarno Seppänen

In this paper, we present exploitability descent, a new algorithm to compute approximate equilibria in two-player zero-sum extensive-form games with imperfect information, by direct policy optimization against worst-case opponents. We prove…

AI in Math deals with mathematics in a constructive manner so that reasoning becomes automated, less laborious, and less error-prone. For algorithms, the question becomes how to automate analyses for specific problems. For the first time,…

计算机科学与博弈论 · 计算机科学 2023-10-13 Xiaotie Deng , Dongchen Li , Hanyu Li

In this paper, we consider a distributed learning problem in a subnetwork zero-sum game, where agents are competing in different subnetworks. These agents are connected through time-varying graphs where each agent has its own cost function…

最优化与控制 · 数学 2021-08-05 Shijie Huang , Jinlong Lei , Yiguang Hong , Uday V. Shanbhag , Jie Chen

We study a target coverage problem in which a team of sensing agents, operating under limited communication, must collaboratively monitor targets that may be adaptively repositioned by an attacker. We model this interaction as a zero-sum…

系统与控制 · 电气工程与系统科学 2026-03-19 Jayanth Bhargav , Zirui Xu , Vasileios Tzoumas , Mahsa Ghasemi , Shreyas Sundaram

We investigate optimal decision making under imperfect recall, that is, when an agent forgets information it once held before. An example is the absentminded driver game, as well as team games in which the members have limited communication…

计算机科学与博弈论 · 计算机科学 2024-06-25 Emanuel Tewolde , Brian Hu Zhang , Caspar Oesterheld , Manolis Zampetakis , Tuomas Sandholm , Paul W. Goldberg , Vincent Conitzer

In this paper, we examine the Nash equilibrium convergence properties of no-regret learning in general N-player games. For concreteness, we focus on the archetypal follow the regularized leader (FTRL) family of algorithms, and we consider…

计算机科学与博弈论 · 计算机科学 2021-02-05 Angeliki Giannou , Emmanouil-Vasileios Vlatakis-Gkaragkounis , Panayotis Mertikopoulos

The current state of the art in playing many important perfect information games, including Chess and Go, combines planning and deep reinforcement learning with self-play. We extend this approach to imperfect information games and present…

人工智能 · 计算机科学 2018-10-26 Andy Kitchen , Michela Benedetti

As quantum processors advance, the emergence of large-scale decentralized systems involving interacting quantum-enabled agents is on the horizon. Recent research efforts have explored quantum versions of Nash and correlated equilibria as…

计算机科学与博弈论 · 计算机科学 2024-12-18 Wayne Lin , Georgios Piliouras , Ryann Sim , Antonios Varvitsiotis

High-quality information set abstraction remains a core challenge in solving large-scale imperfect-information extensive-form games (IIEFGs)--such as no-limit Texas Hold'em--where the finite nature of spatial resources hinders solving…

人工智能 · 计算机科学 2025-12-10 Yanchang Fu , Shengda Liu , Pei Xu , Kaiqi Huang

We provide a formal definition of depth-limited games together with an accessible and rigorous explanation of the underlying concepts, both of which were previously missing in imperfect-information games. The definition works for an…

人工智能 · 计算机科学 2022-03-25 Vojtěch Kovařík , Dominik Seitz , Viliam Lisý , Jan Rudolf , Shuo Sun , Karel Ha

We formulate and analyze a general class of stochastic dynamic games with asymmetric information arising in dynamic systems. In such games, multiple strategic agents control the system dynamics and have different information about the…

计算机科学与博弈论 · 计算机科学 2015-10-26 Yi Ouyang , Hamidreza Tavafoghi , Demosthenis Teneketzis

Traditional methods for computing equilibria in auctions become computationally intractable as auction complexity increases, particularly in multi-item and dynamic auctions. This paper introduces a self-play based reinforcement learning…

综合经济学 · 经济学 2024-10-21 Pranjal Rawat

We consider a game-theoretic model of information retrieval with strategic authors. We examine two different utility schemes: authors who aim at maximizing exposure and authors who want to maximize active selection of their content (i.e.…

计算机科学与博弈论 · 计算机科学 2019-02-21 Omer Ben-Porat , Itay Rosenberg , Moshe Tennenholtz