中文
相关论文

相关论文: A Generalized Minimax Q-learning Algorithm for Two…

200 篇论文

MinMaxMin $Q$-learning is a novel optimistic Actor-Critic algorithm that addresses the problem of overestimation bias ($Q$-estimations are overestimating the real $Q$-values) inherent in conservative RL algorithms. Its core formula relies…

机器学习 · 计算机科学 2024-06-04 Nitsan Soffair , Shie Mannor

Despite the success of generative adversarial networks (GANs) in generating visually appealing images, they are notoriously challenging to train. In order to stabilize the learning dynamics in minimax games, we propose a novel recursive…

机器学习 · 计算机科学 2022-11-01 Zichu Liu , Lacra Pavel

We present a systematic investigation of the quantum games, constructed using a novel repeated game protocol, when played repeatedly ad infinitum. We focus on establishing that such repeated games -- by virtue of inherent quantum-mechanical…

量子物理 · 物理学 2024-02-27 Archan Mukhopadhyay , Saikat Sur , Tanay Saha , Shubhadeep Sadhukhan , Sagar Chakraborty

We study two-player general sum repeated finite games where the rewards of each player are generated from an unknown distribution. Our aim is to find the egalitarian bargaining solution (EBS) for the repeated game, which can lead to much…

机器学习 · 计算机科学 2019-06-05 Aristide Tossou , Christos Dimitrakakis , Jaroslaw Rzepecki , Katja Hofmann

A seminal result in game theory is von Neumann's minmax theorem, which states that zero-sum games admit an essentially unique equilibrium solution. Classical learning results build on this theorem to show that online no-regret dynamics…

计算机科学与博弈论 · 计算机科学 2021-11-08 Tanner Fiez , Ryann Sim , Stratis Skoulakis , Georgios Piliouras , Lillian Ratliff

This paper considers a two-player game where each player chooses a resource from a finite collection of options. Each resource brings a random reward. Both players have statistical information regarding the rewards of each resource.…

计算机科学与博弈论 · 计算机科学 2023-09-19 Mevan Wijewardena , Michael J. Neely

We study a multi-agent reinforcement learning dynamics, and analyze its asymptotic behavior in infinite-horizon discounted Markov potential games. We focus on the independent and decentralized setting, where players do not know the game…

机器学习 · 计算机科学 2025-04-02 Chinmay Maheshwari , Manxi Wu , Druv Pai , Shankar Sastry

Simple stochastic games are two-player zero-sum stochastic games with turn-based moves, perfect information, and reachability winning conditions. We present two new algorithms computing the values of simple stochastic games. Both of them…

计算机科学与博弈论 · 计算机科学 2015-07-01 Hugo Gimbert , Florian Horn

We propose a novel independent and payoff-based learning framework for stochastic games that is model-free, game-agnostic, and gradient-free. The learning dynamics follow a best-response-type actor-critic architecture, where agents update…

机器学习 · 计算机科学 2026-02-03 Ahmed Said Donmez , Yuksel Arslantas , Muhammed O. Sayin

We study two-player zero-sum concurrent stochastic games with finite state and action space played for an infinite number of steps. In every step, the two players simultaneously and independently choose an action. Given the current state…

计算机科学与博弈论 · 计算机科学 2024-10-10 Ali Asadi , Krishnendu Chatterjee , Raimundo Saona , Jakub Svoboda

In this paper, we study the problem of learning in quantum games - and other classes of semidefinite games - with scalar, payoff-based feedback. For concreteness, we focus on the widely used matrix multiplicative weights (MMW) algorithm…

计算机科学与博弈论 · 计算机科学 2023-11-07 Kyriakos Lotidis , Panayotis Mertikopoulos , Nicholas Bambos , Jose Blanchet

We provide several applications of Optimistic Mirror Descent, an online learning algorithm based on the idea of predictable sequences. First, we recover the Mirror Prox algorithm for offline optimization, prove an extension to Holder-smooth…

机器学习 · 计算机科学 2013-11-11 Alexander Rakhlin , Karthik Sridharan

In this paper, we study inverse game theory (resp. inverse multiagent learning) in which the goal is to find parameters of a game's payoff functions for which the expected (resp. sampled) behavior is an equilibrium. We formulate these…

计算机科学与博弈论 · 计算机科学 2025-02-21 Denizalp Goktas , Amy Greenwald , Sadie Zhao , Alec Koppel , Sumitra Ganesh

With the recent advances in solving large, zero-sum extensive form games, there is a growing interest in the inverse problem of inferring underlying game parameters given only access to agent actions. Although a recent work provides a…

机器学习 · 计算机科学 2019-03-12 Chun Kai Ling , Fei Fang , J. Zico Kolter

We devise a policy-iteration algorithm for deterministic two-player discounted and mean-payoff games, that runs in polynomial time with high probability, on any input where each payoff is chosen independently from a sufficiently random…

计算机科学与博弈论 · 计算机科学 2024-02-07 Bruno Loff , Mateusz Skomra

We consider a two-player zero-sum stochastic differential game in which one of the players has a private information on the game. Both players observe each other, so that the non-informed player can try to guess his missing information. Our…

概率论 · 数学 2011-06-15 Christine Grün

This article investigates the optimal control problem with disturbance rejection for discrete-time multi-agent systems under cooperative and non-cooperative graphical games frameworks. Given the practical challenges of obtaining accurate…

系统与控制 · 电气工程与系统科学 2025-04-11 Xinyang Wang , Martin Guay , Shimin Wang , Hongwei Zhang

Many tasks in modern machine learning can be formulated as finding equilibria in \emph{sequential} games. In particular, two-player zero-sum sequential games, also known as minimax optimization, have received growing interest. It is…

机器学习 · 计算机科学 2019-11-26 Yuanhao Wang , Guodong Zhang , Jimmy Ba

The paper is devoted to the first-order mean field game system in the case when the distribution of players can contain atoms. The proposed definition of a generalized solution is based on the minimax approach to the Hamilton-Jacobi…

偏微分方程分析 · 数学 2014-04-21 Yurii Averboukh

In this paper, the problem of false information injection attack and defense on state estimation in dynamic multi-sensor systems is investigated from a game theoretic perspective. The relationship between the Kalman filter and the adversary…

系统与控制 · 计算机科学 2015-02-13 Jingyang Lu , Ruixin Niu
‹ 上一页 1 8 9 10 下一页 ›