中文
相关论文

相关论文: Mutation-Driven Follow the Regularized Leader for …

200 篇论文

Follow the regularized leader FTRL is the premier algorithm for online optimization. However, despite decades of research on its convergence in constrained optimization -- and potential games in particular -- its behavior remained hitherto…

计算机科学与博弈论 · 计算机科学 2026-02-02 Ioannis Anagnostides , Ioannis Panageas , Nikolas Patris , Tuomas Sandholm

Understanding the behavior of no-regret dynamics in general $N$-player games is a fundamental question in online learning and game theory. A folk result in the field states that, in finite games, the empirical frequency of play under…

We prove that optimistic-follow-the-regularized-leader (OFTRL), together with smooth value updates, finds an $O(T^{-1})$-approximate Nash equilibrium in $T$ iterations for two-player zero-sum Markov games with full information. This…

机器学习 · 计算机科学 2023-02-10 Yuepeng Yang , Cong Ma

Follow-the-regularized-leader (FTRL) algorithms have become popular in the context of games, providing easy-to-implement methods for each agent, as well as theoretical guarantees that the strategies of all agents will converge to some…

系统与控制 · 电气工程与系统科学 2026-03-31 Heling Zhang , Siqi Du , Roy Dong

In this paper, we investigate how randomness and uncertainty influence learning in games. Specifically, we examine a perturbed variant of the dynamics of "follow-the-regularized-leader" (FTRL), where the players' payoff observations and…

计算机科学与博弈论 · 计算机科学 2025-06-17 Pierre-Louis Cauvin , Davide Legacci , Panayotis Mertikopoulos

In this paper, we examine the Nash equilibrium convergence properties of no-regret learning in general N-player games. For concreteness, we focus on the archetypal follow the regularized leader (FTRL) family of algorithms, and we consider…

计算机科学与博弈论 · 计算机科学 2021-02-05 Angeliki Giannou , Emmanouil-Vasileios Vlatakis-Gkaragkounis , Panayotis Mertikopoulos

Bargaining games, where agents attempt to agree on how to split utility, are an important class of games used to study economic behavior, which motivates a study of online learning algorithms in these games. In this work, we tackle when…

计算机科学与博弈论 · 计算机科学 2025-07-08 Serafina Kamp , Reese Liebman , Benjamin Fish

We investigate how perturbation does and does not improve the Follow-the-Regularized-Leader (FTRL) algorithm in solving imperfect-information extensive-form games under sampling, where payoffs are estimated from sampled trajectories. While…

计算机科学与博弈论 · 计算机科学 2025-08-05 Wataru Masaka , Mitsuki Sakamoto , Kenshi Abe , Kaito Ariu , Tuomas Sandholm , Atsushi Iwasaki

The long-run behavior of multi-agent learning - and, in particular, no-regret learning - is relatively well-understood in potential games, where players have aligned interests. By contrast, in harmonic games - the strategic counterpart of…

计算机科学与博弈论 · 计算机科学 2024-12-31 Davide Legacci , Panayotis Mertikopoulos , Christos H. Papadimitriou , Georgios Piliouras , Bary S. R. Pradelski

We investigate the accuracy of prediction in deterministic learning dynamics of zero-sum games with random initializations, specifically focusing on observer uncertainty and its relationship to the evolution of covariances. Zero-sum games…

计算机科学与博弈论 · 计算机科学 2024-06-18 Yi Feng , Georgios Piliouras , Xiao Wang

This paper investigates the impact of feedback quantization on multi-agent learning. In particular, we analyze the equilibrium convergence properties of the well-known "follow the regularized leader" (FTRL) class of algorithms when players…

计算机科学与博弈论 · 计算机科学 2022-09-13 Kyriakos Lotidis , Panayotis Mertikopoulos , Nicholas Bambos

We derive a new analysis of Follow The Regularized Leader (FTRL) for online learning with delayed bandit feedback. By separating the cost of delayed feedback from that of bandit feedback, our analysis allows us to obtain new results in…

机器学习 · 计算机科学 2023-05-16 Dirk van der Hoeven , Lukas Zierahn , Tal Lancewicki , Aviv Rosenberg , Nicoló Cesa-Bianchi

The linear bandit problem has been studied for many years in both stochastic and adversarial settings. Designing an algorithm that can optimize the environment without knowing the loss type attracts lots of interest. \citet{LeeLWZ021}…

机器学习 · 计算机科学 2023-07-19 Fang Kong , Canzhe Zhao , Shuai Li

We consider a common case of the combinatorial semi-bandit problem, the $m$-set semi-bandit, where the learner exactly selects $m$ arms from the total $d$ arms. In the adversarial setting, the best regret bound, known to be…

机器学习 · 计算机科学 2025-07-08 Jingxin Zhan , Yuchen Xin , Chenjie Sun , Zhihua Zhang

In this paper we extend the classical Follow-The-Regularized-Leader (FTRL) algorithm to encompass time-varying constraints, through adaptive penalization. We establish sufficient conditions for the proposed Penalized FTRL algorithm to…

机器学习 · 计算机科学 2022-04-07 Douglas J. Leith , George Iosifidis

This study raises and addresses the problem of time-delayed feedback in learning in games. Because learning in games assumes that multiple agents independently learn their strategies, a discrepancy in optimization often emerges among the…

机器学习 · 计算机科学 2025-11-10 Yuma Fujimoto , Kenshi Abe , Kaito Ariu

Zero-sum games are a fundamental setting for adversarial training and decision-making in multi-agent learning (MAL). Existing methods often ensure convergence to (approximate) Nash equilibria by introducing a form of regularization. Yet,…

多智能体系统 · 计算机科学 2026-02-10 Tuo Zhang , Leonardo Stella

This paper studies the optimistic variant of Fictitious Play for learning in two-player zero-sum games. While it is known that Optimistic FTRL -- a regularized algorithm with a bounded stepsize parameter -- obtains constant regret in this…

机器学习 · 计算机科学 2026-01-15 John Lazarsfeld , Georgios Piliouras , Ryann Sim , Stratis Skoulakis

Regret minimization is a powerful method for finding Nash equilibria in Normal-Form Games (NFGs) and Extensive-Form Games (EFGs), but it typically guarantees convergence only for the average strategy. However, computing the average strategy…

计算机科学与博弈论 · 计算机科学 2025-09-18 Hang Ren , Yulin Wu , Shuhan Qi , Jiajia Zhang , Xiaozhen Sun , Tianzi Ma , Xuan Wang

We study the problem of designing adaptive multi-armed bandit algorithms that perform optimally in both the stochastic setting and the adversarial setting simultaneously (often known as a best-of-both-world guarantee). A line of recent…

机器学习 · 计算机科学 2023-10-27 Tiancheng Jin , Junyan Liu , Haipeng Luo
‹ 上一页 1 2 3 10 下一页 ›