中文
相关论文

相关论文: Online Double Oracle

200 篇论文

By incorporating regret minimization, double oracle methods have demonstrated rapid convergence to Nash Equilibrium (NE) in normal-form games and extensive-form games, through algorithms such as online double oracle (ODO) and extensive-form…

计算机科学与博弈论 · 计算机科学 2023-07-14 Xiaohang Tang , Le Cong Dinh , Stephen Marcus McAleer , Yaodong Yang

Policy Space Response Oracles (PSRO) is a reinforcement learning (RL) algorithm for two-player zero-sum games that has been empirically shown to find approximate Nash equilibria in large games. Although PSRO is guaranteed to converge to an…

计算机科学与博弈论 · 计算机科学 2022-02-01 Stephen McAleer , John Lanier , Kevin Wang , Pierre Baldi , Roy Fox

We study a novel setting in Online Markov Decision Processes (OMDPs) where the loss function is chosen by a non-oblivious strategic adversary who follows a no-external regret algorithm. In this setting, we first demonstrate that MDP-Expert,…

机器学习 · 计算机科学 2023-01-31 Le Cong Dinh , David Henry Mguni , Long Tran-Thanh , Jun Wang , Yaodong Yang

Policy space response oracles (PSRO) is a multi-agent reinforcement learning algorithm that has achieved state-of-the-art performance in very large two-player zero-sum games. PSRO is based on the tabular double oracle (DO) method, an…

计算机科学与博弈论 · 计算机科学 2022-02-01 Stephen McAleer , Kevin Wang , John Lanier , Marc Lanctot , Pierre Baldi , Tuomas Sandholm , Roy Fox

Many efficient algorithms have been designed to recover Nash equilibria of various classes of finite games. Special classes of continuous games with infinite strategy spaces, such as polynomial games, can be solved by semidefinite…

计算机科学与博弈论 · 计算机科学 2020-10-01 Lukáš Adam , Rostislav Horčík , Tomáš Kasl , Tomáš Kroupa

Policy Space Response Oracle methods (PSRO) provide a general solution to learn Nash equilibrium in two-player zero-sum games but suffer from two drawbacks: (1) the computation inefficiency due to the need for consistent meta-game…

计算机科学与博弈论 · 计算机科学 2022-06-02 Ming Zhou , Jingxiao Chen , Ying Wen , Weinan Zhang , Yaodong Yang , Yong Yu , Jun Wang

Extensive-Form Game (EFG) represents a fundamental model for analyzing sequential interactions among multiple agents and the primary challenge to solve it lies in mitigating sample complexity. Existing research indicated that Double Oracle…

计算机科学与博弈论 · 计算机科学 2024-11-05 Xiaohang Tang , Chiyuan Wang , Chengdong Ma , Ilija Bogunovic , Stephen McAleer , Yaodong Yang

We consider online learning in multi-player smooth monotone games. Existing algorithms have limitations such as (1) being only applicable to strongly monotone games; (2) lacking the no-regret guarantee; (3) having only asymptotic or slow…

机器学习 · 计算机科学 2023-09-06 Yang Cai , Weiqiang Zheng

While ERM suffices to attain near-optimal generalization error in the stochastic learning setting, this is not known to be the case in the online learning setting, where algorithms for general concept classes rely on computationally…

机器学习 · 计算机科学 2023-07-11 Angelos Assos , Idan Attias , Yuval Dagan , Constantinos Daskalakis , Maxwell Fishelson

Motivated by alternating learning dynamics in two-player games, a recent work by Cevher et al.(2024) shows that $o(\sqrt{T})$ alternating regret is possible for any $T$-round adversarial Online Linear Optimization (OLO) problem, and left as…

机器学习 · 计算机科学 2025-06-19 Soumita Hait , Ping Li , Haipeng Luo , Mengxiao Zhang

Centered around solving the Online Saddle Point problem, this paper introduces the Online Convex-Concave Optimization (OCCO) framework, which involves a sequence of two-player time-varying convex-concave games. We propose the generalized…

机器学习 · 计算机科学 2023-12-18 Qing-xin Meng , Jian-wei Liu

This paper investigates a population-based training regime based on game-theoretic principles called Policy-Spaced Response Oracles (PSRO). PSRO is general in the sense that it (1) encompasses well-known algorithms such as fictitious play…

We propose the first online quantum algorithm for solving zero-sum games with $\widetilde O(1)$ regret under the game setting. Moreover, our quantum algorithm computes an $\varepsilon$-approximate Nash equilibrium of an $m \times n$ matrix…

量子物理 · 物理学 2024-10-01 Minbo Gao , Zhengfeng Ji , Tongyang Li , Qisheng Wang

We study online learning in two-player uninformed Markov games, where the opponent's actions and policies are unobserved. In this setting, Tian et al. (2021) show that achieving no-external-regret is impossible without incurring an…

机器学习 · 计算机科学 2026-02-10 Junyan Liu , Haipeng Luo , Zihan Zhang , Lillian J. Ratliff

We study the problem of game-theoretic robot allocation where two players strategically allocate robots to compete for multiple sites of interest. Robots possess offensive or defensive capabilities to interfere and weaken their opponents to…

机器人学 · 计算机科学 2025-03-25 Zijian An , Lifeng Zhou

Offline learning of strategies takes data efficiency to its extreme by restricting algorithms to a fixed dataset of state-action trajectories. We consider the problem in a mixed-motive multiagent setting, where the goal is to solve a game…

人工智能 · 计算机科学 2026-03-03 Austin A. Nguyen , Michael P. Wellman

In competitive two-agent environments, deep reinforcement learning (RL) methods based on the \emph{Double Oracle (DO)} algorithm, such as \emph{Policy Space Response Oracles (PSRO)} and \emph{Anytime PSRO (APSRO)}, iteratively add RL best…

计算机科学与博弈论 · 计算机科学 2022-07-15 Stephen McAleer , JB Lanier , Kevin Wang , Pierre Baldi , Roy Fox , Tuomas Sandholm

Self-play via online learning is one of the premier ways to solve large-scale two-player zero-sum games, both in theory and practice. Particularly popular algorithms include optimistic multiplicative weights update (OMWU) and optimistic…

计算机科学与博弈论 · 计算机科学 2025-01-22 Yang Cai , Gabriele Farina , Julien Grand-Clément , Christian Kroer , Chung-Wei Lee , Haipeng Luo , Weiqiang Zheng

Reinforcement learning (RL) in large environments often suffers from severe computational bottlenecks, as conventional regret minimization algorithms require repeated, costly calls to planning and statistical estimation oracles. While…

机器学习 · 计算机科学 2026-05-04 Haichen Hu , Jian Qian , David Simchi-Levi

No-regret learning has been widely used to compute a Nash equilibrium in two-person zero-sum games. However, there is still a lack of regret analysis for network stochastic zero-sum games, where players competing in two subnetworks only…

最优化与控制 · 数学 2022-05-31 Shijie Huang , Jinlong Lei , Yiguang Hong
‹ 上一页 1 2 3 10 下一页 ›