中文
相关论文

相关论文: Game Redesign in No-regret Game Playing

200 篇论文

The problem of matching markets has been studied for a long time in the literature due to its wide range of applications. Finding a stable matching is a common equilibrium objective in this problem. Since market participants are usually…

机器学习 · 计算机科学 2023-07-21 Fang Kong , Shuai Li

Scale-invariance in games has recently emerged as a widely valued desirable property. Yet, almost all fast convergence guarantees in learning in games require prior knowledge of the utility scale. To address this, we develop learning…

计算机科学与博弈论 · 计算机科学 2026-02-13 Taira Tsuchiya , Haipeng Luo , Shinji Ito

We consider the problem of online learning where the sequence of actions played by the learner must adhere to an unknown safety constraint at every round. The goal is to minimize regret with respect to the best safe action in hindsight…

机器学习 · 计算机科学 2024-03-08 Karthik Sridharan , Seung Won Wilson Yoo

We provide a novel reduction from swap-regret minimization to external-regret minimization, which improves upon the classical reductions of Blum-Mansour [BM07] and Stolz-Lugosi [SL05] in that it does not require finiteness of the space of…

机器学习 · 计算机科学 2025-02-25 Yuval Dagan , Constantinos Daskalakis , Maxwell Fishelson , Noah Golowich

No-regret learning has a long history of being closely connected to game theory. Recent works have devised uncoupled no-regret learning dynamics that, when adopted by all the players in normal-form games, converge to various equilibrium…

计算机科学与博弈论 · 计算机科学 2024-04-24 Weichao Mao , Haoran Qiu , Chen Wang , Hubertus Franke , Zbigniew Kalbarczyk , Tamer Başar

Motivated by the stringent safety requirements that are often present in real-world applications, we study a safe online convex optimization setting where the player needs to simultaneously achieve sublinear regret and zero constraint…

机器学习 · 计算机科学 2024-07-17 Spencer Hutchinson , Mahnoosh Alizadeh

This work studies external regret in sequential prediction games with both positive and negative payoffs. External regret measures the difference between the payoff obtained by the forecasting strategy and the payoff of the best action. In…

统计理论 · 数学 2007-06-13 Nicolo Cesa-Bianchi , Yishay Mansour , Gilles Stoltz

In games with a large number of players where players may have overlapping objectives, the analysis of stable outcomes typically depends on player types. A special case is when a large part of the player population consists of imitation…

计算机科学与博弈论 · 计算机科学 2010-06-18 Soumya Paul , R. Ramanujam

We present an algorithm which attains O(\sqrt{T}) internal (and thus external) regret for finite games with partial monitoring under the local observability condition. Recently, this condition has been shown by (Bartok, Pal, and Szepesvari,…

机器学习 · 计算机科学 2011-09-01 Dean Foster , Alexander Rakhlin

We study the power of different types of adaptive (nonoblivious) adversaries in the setting of prediction with expert advice, under both full-information and bandit feedback. We measure the player's performance using a new notion of regret,…

机器学习 · 计算机科学 2013-06-04 Nicolo Cesa-Bianchi , Ofer Dekel , Ohad Shamir

We consider an online regression setting in which individuals adapt to the regression model: arriving individuals are aware of the current model, and invest strategically in modifying their own features so as to improve the predicted score…

机器学习 · 计算机科学 2021-03-02 Yahav Bechavod , Katrina Ligett , Zhiwei Steven Wu , Juba Ziani

A rich class of mechanism design problems can be understood as incomplete-information games between a principal who commits to a policy and an agent who responds, with payoffs determined by an unknown state of the world. Traditionally,…

理论经济学 · 经济学 2020-09-14 Modibo Camara , Jason Hartline , Aleck Johnsen

An urban planner might design the spatial layout of transportation amenities so as to improve accessibility for underserved communities -- a fairness objective. However, implementing such a design might trigger processes of neighborhood…

多智能体系统 · 计算机科学 2024-02-12 J. Carlos Martínez Mori , Zhanzhan Zhao

Reinforcement-learning agents seek to maximize a reward signal through environmental interactions. As humans, our job in the learning process is to design reward functions to express desired behavior and enable the agent to learn such…

机器学习 · 计算机科学 2024-08-08 Zhiyuan Zhou , Shreyas Sundara Raman , Henry Sowerby , Michael L. Littman

Publishers who publish their content on the web act strategically, in a behavior that can be modeled within the online learning framework. Regret, a central concept in machine learning, serves as a canonical measure for assessing the…

计算机科学与博弈论 · 计算机科学 2025-01-30 Omer Madmon , Idan Pipano , Itamar Reinman , Moshe Tennenholtz

Algorithmic recourse provides individuals who receive undesirable outcomes from machine learning systems with minimum-cost improvements to achieve a desirable outcome. However, machine learning models often get updated, so the recourse may…

机器学习 · 计算机科学 2026-04-28 Kshitij Kayastha , Vasilis Gkatzelis , Shahin Jabbari

An abundance of recent impossibility results establish that regret minimization in Markov games with adversarial opponents is both statistically and computationally intractable. Nevertheless, none of these results preclude the possibility…

机器学习 · 计算机科学 2025-06-17 Liad Erez , Tal Lancewicki , Uri Sherman , Tomer Koren , Yishay Mansour

Regret minimization is treated as the golden rule in the traditional study of online learning. However, regret minimization algorithms tend to converge to the static optimum, thus being suboptimal for changing environments. To address this…

机器学习 · 计算机科学 2020-02-07 Lijun Zhang , Shiyin Lu , Tianbao Yang

To regulate a social system comprised of self-interested agents, economic incentives are often required to induce a desirable outcome. This incentive design problem naturally possesses a bilevel structure, in which a designer modifies the…

计算机科学与博弈论 · 计算机科学 2022-10-14 Boyi Liu , Jiayang Li , Zhuoran Yang , Hoi-To Wai , Mingyi Hong , Yu Marco Nie , Zhaoran Wang

The connection between games and no-regret algorithms has been widely studied in the literature. A fundamental result is that when all players play no-regret strategies, this produces a sequence of actions whose time-average is a…

计算机科学与博弈论 · 计算机科学 2020-09-15 Zhe Feng , Guru Guruganesh , Christopher Liaw , Aranyak Mehta , Abhishek Sethi