中文
相关论文

相关论文: Winning Without Observing Payoffs: Exploiting Beha…

200 篇论文

We study a sequential coin-flipping game in which a player starts with~$n$ coins, each landing heads independently with probability~$p$. In each round the player flips all remaining coins and must set aside at least one coin showing heads;…

概率论 · 数学 2026-04-28 Peter Pfaffelhuber

We extend the optimin notion of Ismail (2025) from mixed strategy profiles to correlated distributions. A correlated distribution is evaluated by the worst expected payoff each player can receive when opponents may either obey their private…

理论经济学 · 经济学 2026-05-20 Mehmet Mars Seven

We introduce a simple but general online learning framework in which a learner plays against an adversary in a vector-valued game that changes every round. Even though the learner's objective is not convex-concave (and so the minimax…

机器学习 · 计算机科学 2022-10-14 Daniel Lee , Georgy Noarov , Mallesh Pai , Aaron Roth

A recent body of experimental literature has studied empirical game-theoretical analysis, in which we have partial knowledge of a game, consisting of observations of a subset of the pure-strategy profiles and their associated payoffs to…

计算机科学与博弈论 · 计算机科学 2014-02-13 John Fearnley , Martin Gairing , Paul Goldberg , Rahul Savani

We derive asymptotically optimal statistical decision rules for discrete choice problems when payoffs depend on a partially-identified parameter $\theta$ and the decision maker can use a point-identified parameter $\mu$ to deduce…

计量经济学 · 经济学 2025-12-19 Timothy Christensen , Hyungsik Roger Moon , Frank Schorfheide

We consider a model where an agent has a repeated decision to make and wishes to maximize their total payoff. Payoffs are influenced by an action taken by the agent, but also an unknown state of the world that evolves over time. Before…

计算机科学与博弈论 · 计算机科学 2021-01-20 Nicole Immorlica , Ian Kash , Brendan Lucier

We study zero-sum differential games with state constraints and one-sided information, where the informed player (Player 1) has a categorical payoff type unknown to the uninformed player (Player 2). The goal of Player 1 is to minimize his…

计算机科学与博弈论 · 计算机科学 2024-06-05 Mukesh Ghimire , Lei Zhang , Zhe Xu , Yi Ren

We present Self-Play Preference Optimization (SPO), an algorithm for reinforcement learning from human feedback. Our approach is minimalist in that it does not require training a reward model nor unstable adversarial training and is…

机器学习 · 计算机科学 2024-06-14 Gokul Swamy , Christoph Dann , Rahul Kidambi , Zhiwei Steven Wu , Alekh Agarwal

We study a version of the classical zero-sum matrix game with unknown payoff matrix and bandit feedback, where the players only observe each others actions and a noisy payoff. This generalizes the usual matrix game, where the payoff matrix…

机器学习 · 计算机科学 2021-06-15 Brendan O'Donoghue , Tor Lattimore , Ian Osband

In a zero-sum stochastic game, at each stage, two adversary players take decisions and receive a stage payoff determined by them and by a controlled random variable representing the state of nature. The total payoff is the normalized…

最优化与控制 · 数学 2022-05-06 Olivier Catoni , Miquel Oliu-Barton , Bruno Ziliotto

We study players interacting under the veil of ignorance, who have -- coarse -- beliefs represented as subsets of opponents' actions. We analyze when these players follow $\max \min$ or $\max\max$ decision criteria, which we identify with…

理论经济学 · 经济学 2022-11-10 Pierfrancesco Guarino , Gabriel Ziegler

We study the problem of characterizing optimal learning algorithms for playing repeated games against an adversary with unknown payoffs. In this problem, the first player (called the learner) commits to a learning algorithm against a second…

计算机科学与博弈论 · 计算机科学 2024-02-16 Eshwar Ram Arunachaleswaran , Natalie Collina , Jon Schneider

We consider how an agent should update her uncertainty when it is represented by a set P of probability distributions and the agent observes that a random variable X takes on value x, given that the agent makes decisions using the minimax…

人工智能 · 计算机科学 2014-07-29 Peter D. Grunwald , Joseph Y. Halpern

In this article, we focus on search algorithms for two-player perfect information games, whose objective is to determine the best possible strategy, and ideally a winning strategy. Unfortunately, some search algorithms for games in the…

人工智能 · 计算机科学 2026-03-26 Quentin Cohen-Solal

We consider multiplayer stochastic games in which the payoff of each player is a bounded and Borel-measurable function of the infinite play. By using a generalization of the technique of Martin (1998) and Maitra and Sudderth (1998), we show…

最优化与控制 · 数学 2022-08-26 János Flesch , Eilon Solan

Game theory has grown into a major field over the past few decades, and poker has long served as one of its key case studies. Game-Theory-Optimal (GTO) provides strategies to avoid loss in poker, but pure GTO does not guarantee maximum…

计算机科学与博弈论 · 计算机科学 2025-09-30 SeungHyun Yi , Seungjun Yi

We consider online learning when the time horizon is unknown. We apply a minimax analysis, beginning with the fixed horizon case, and then moving on to two unknown-horizon settings, one that assumes the horizon is chosen randomly according…

机器学习 · 计算机科学 2013-10-08 Haipeng Luo , Robert E. Schapire

We introduce the study of search games between a mobile Searcher and an immobile Hider in a new setting in which the Searcher has some potentially erroneous information, i.e., a prediction on the Hider's position. The objective is to…

计算机科学与博弈论 · 计算机科学 2024-09-05 Spyros Angelopoulos , Thomas Lidbetter , Konstantinos Panagiotou

We consider how an agent should update her uncertainty when it is represented by a set $\P$ of probability distributions and the agent observes that a random variable $X$ takes on value $x$, given that the agent makes decisions using the…

人工智能 · 计算机科学 2007-11-27 Peter D. Grunwald , Joseph Y. Halpern

In an adversarial environment, a hostile player performing a task may behave like a non-hostile one in order not to reveal its identity to an opponent. To model such a scenario, we define identity concealment games: zero-sum stochastic…

计算机科学与博弈论 · 计算机科学 2024-03-05 Mustafa O. Karabag , Melkior Ornik , Ufuk Topcu