中文
相关论文

相关论文: Near-Optimal No-Regret Learning Dynamics for Gener…

200 篇论文

In this paper, we consider a distributed learning problem in a subnetwork zero-sum game, where agents are competing in different subnetworks. These agents are connected through time-varying graphs where each agent has its own cost function…

最优化与控制 · 数学 2021-08-05 Shijie Huang , Jinlong Lei , Yiguang Hong , Uday V. Shanbhag , Jie Chen

This paper addresses safe distributed online optimization over an unknown set of linear safety constraints. A network of agents aims at jointly minimizing a global, time-varying function, which is only partially observable to each…

最优化与控制 · 数学 2023-02-27 Ting-Jui Chang , Sapana Chaudhary , Dileep Kalathil , Shahin Shahrampour

We give a simple and computationally efficient algorithm that, for any constant $\varepsilon>0$, obtains $\varepsilon T$-swap regret within only $T = \mathsf{polylog}(n)$ rounds; this is an exponential improvement compared to the…

计算机科学与博弈论 · 计算机科学 2023-11-15 Binghui Peng , Aviad Rubinstein

Regret-based algorithms are highly efficient at finding approximate Nash equilibria in sequential games such as poker games. However, most regret-based algorithms, including counterfactual regret minimization (CFR) and its variants, rely on…

机器学习 · 计算机科学 2021-10-28 Chung-Wei Lee , Christian Kroer , Haipeng Luo

This paper examines the convergence of no-regret learning in games with continuous action sets. For concreteness, we focus on learning via "dual averaging", a widely used class of no-regret learning schemes where players take small steps…

最优化与控制 · 数学 2018-01-17 Panayotis Mertikopoulos , Zhengyuan Zhou

Projection-free online learning has drawn increasing interest due to its efficiency in solving high-dimensional problems with complicated constraints. However, most existing projection-free online methods focus on minimizing the static…

机器学习 · 计算机科学 2023-05-22 Yibo Wang , Wenhao Yang , Wei Jiang , Shiyin Lu , Bing Wang , Haihong Tang , Yuanyu Wan , Lijun Zhang

This paper introduces the new concept of (follower) satisfaction in Stackelberg games and compares the standard Stackelberg game with its satisfaction version. Simulation results are presented which suggest that the follower adopting…

计算机科学与博弈论 · 计算机科学 2024-08-22 Langford White , Duong Nguyen , Hung Nguyen

Linear bandit algorithms yield $\tilde{\mathcal{O}}(n\sqrt{T})$ pseudo-regret bounds on compact convex action sets $\mathcal{K}\subset\mathbb{R}^n$ and two types of structural assumptions lead to better pseudo-regret bounds. When…

机器学习 · 计算机科学 2021-03-11 Thomas Kerdreux , Christophe Roux , Alexandre d'Aspremont , Sebastian Pokutta

Constrained Online Convex Optimization (COCO) can be seen as a generalization of the standard Online Convex Optimization (OCO) framework. At each round, a cost function and constraint function are revealed after a learner chooses an action.…

机器学习 · 计算机科学 2025-05-30 Ricardo N. Ferreira , Cláudia Soares

This paper examines the convergence of no-regret learning in Cournot games with continuous actions. Cournot games are the essential model for many socio-economic systems, where players compete by strategically setting their output quantity.…

计算机科学与博弈论 · 计算机科学 2020-02-12 Yuanyuan Shi , Baosen Zhang

Policy optimization methods are popular reinforcement learning algorithms in practice. Recent works have built theoretical foundation for them by proving $\sqrt{T}$ regret bounds even when the losses are adversarial. Such bounds are tight…

机器学习 · 计算机科学 2023-02-21 Christoph Dann , Chen-Yu Wei , Julian Zimmert

Our paper studies the setting of players using no-regret algorithms in various two-player games. We address whether having stronger regret guarantees or playing against an opponent with weaker regret guarantees yields higher utilities for…

计算机科学与博弈论 · 计算机科学 2026-04-29 R. Xu , E. Yachbes , J. Zhang

We study episodic reinforcement learning with fixed reward and transition functions, but with episode-dependent admissible action sets that are observed at the start of each episode. Performance is measured by cumulative regret against the…

机器学习 · 计算机科学 2026-05-18 Zijun Chen , Zihan Zhang

Self-play methods based on regret minimization have become the state of the art for computing Nash equilibria in large two-players zero-sum extensive-form games. These methods fundamentally rely on the hierarchical structure of the players'…

计算机科学与博弈论 · 计算机科学 2019-10-29 Gabriele Farina , Chun Kai Ling , Fei Fang , Tuomas Sandholm

In online convex optimization, the player aims to minimize regret, or the difference between her loss and that of the best fixed decision in hindsight over the entire repeated game. Algorithms that minimize (standard) regret may converge to…

机器学习 · 计算机科学 2023-02-14 Zhou Lu , Elad Hazan

Self-play, where the algorithm learns by playing against itself without requiring any direct supervision, has become the new weapon in modern Reinforcement Learning (RL) for achieving superhuman performance in practice. However, the…

机器学习 · 计算机科学 2020-07-10 Yu Bai , Chi Jin

We study the problem of \emph{dynamic regret minimization} in $K$-armed Dueling Bandits under non-stationary or time varying preferences. This is an online learning setup where the agent chooses a pair of items at each round and observes…

机器学习 · 计算机科学 2022-06-14 Aadirupa Saha , Shubham Gupta

In this paper, we provide a novel and simple algorithm, Clairvoyant Multiplicative Weights Updates (CMWU) for regret minimization in general games. CMWU effectively corresponds to the standard MWU algorithm but where all agents, when…

计算机科学与博弈论 · 计算机科学 2022-06-30 Georgios Piliouras , Ryann Sim , Stratis Skoulakis

To deal with non-stationary online problems with complex constraints, we investigate the dynamic regret of online Frank-Wolfe (OFW), which is an efficient projection-free algorithm for online convex optimization. It is well-known that in…

机器学习 · 计算机科学 2024-06-25 Yuanyu Wan , Lijun Zhang , Mingli Song

We study the control of an \emph{unknown} linear dynamical system under general convex costs. The objective is minimizing regret vs. the class of disturbance-feedback-controllers, which encompasses all stabilizing…

机器学习 · 计算机科学 2020-10-30 Orestis Plevrakis , Elad Hazan