中文
相关论文

相关论文: Zero-sum turn games using Q-learning: finite compu…

200 篇论文

The exploration-exploitation dilemma has been a central challenge in reinforcement learning (RL) with complex model classes. In this paper, we propose a new algorithm, Monotonic Q-Learning with Upper Confidence Bound (MQL-UCB) for RL with…

机器学习 · 计算机科学 2025-10-06 Heyang Zhao , Jiafan He , Quanquan Gu

In state of the art model-free off-policy deep reinforcement learning, a replay memory is used to store past experience and derive all network updates. Even if both state and action spaces are continuous, the replay memory only holds a…

机器学习 · 计算机科学 2020-07-16 Sabrina Hoppe , Marc Toussaint

Consider a two-player zero-sum stochastic game where the transition function can be embedded in a given feature space. We propose a two-player Q-learning algorithm for approximating the Nash equilibrium strategy via sampling. The algorithm…

机器学习 · 计算机科学 2019-06-04 Zeyu Jia , Lin F. Yang , Mengdi Wang

Stochastic games are a convenient formalism for modelling systems that comprise rational agents competing or collaborating within uncertain environments. Probabilistic model checking techniques for this class of models allow us to formally…

计算机科学中的逻辑 · 计算机科学 2022-11-14 Marta Kwiatkowska , Gethin Norman , David Parker , Gabriel Santos

We consider evolutionary games in which the agent selected for update compares their payoff to q neighbours, rather than a single neighbour as in standard evolutionary game theory. Through studying fixed point stability and fixation times…

物理与社会 · 物理学 2025-01-14 Christopher R. Kitching , Tobias Galla

Conventional language model (LM) safety alignment relies on a reactive, disjoint procedure: attackers exploit a static model, followed by defensive fine-tuning to patch exposed vulnerabilities. This sequential approach creates a mismatch --…

机器学习 · 计算机科学 2025-10-07 Mickel Liu , Liwei Jiang , Yancheng Liang , Simon Shaolei Du , Yejin Choi , Tim Althoff , Natasha Jaques

Q-learning is known as one of the fundamental reinforcement learning (RL) algorithms. Its convergence has been the focus of extensive research over the past several decades. Recently, a new finitetime error bound and analysis for Q-learning…

系统与控制 · 电气工程与系统科学 2024-01-17 Donghwna Lee

This paper investigates the sublinear regret guarantees of two non-no-regret algorithms in zero-sum games: Fictitious Play, and Online Gradient Descent with constant stepsizes. In general adversarial online learning settings, both…

机器学习 · 计算机科学 2025-06-17 John Lazarsfeld , Georgios Piliouras , Ryann Sim , Andre Wibisono

Behavioral experiments on the trust game have shown that trust and trustworthiness are universal among human beings, contradicting the prediction by assuming \emph{Homo economicus} in orthodox Economics. This means some mechanism must be at…

种群与进化 · 定量生物学 2024-12-20 Guozhong Zheng , Jiqiang Zhang , Jing Zhang , Weiran Cai , Li Chen

In this paper, we formulate inverse reinforcement learning (IRL) as an expert-learner interaction whereby the optimal performance intent of an expert or target agent is unknown to a learner agent. The learner observes the states and…

机器学习 · 计算机科学 2023-01-06 Wenqian Xue , Bosen Lian , Jialu Fan , Tianyou Chai , Frank L. Lewis

Zero sum games with risk-sensitive cost criterion are considered with underlying dynamics being given by controlled stochastic differential equations. Under the assumption of geometric stability on the dynamics , we completely characterize…

最优化与控制 · 数学 2018-01-04 Anup Biswas , Subhamay Saha

We develop a continuous-time reinforcement learning framework for a class of singular stochastic control problems without entropy regularization. The optimal singular control is characterized as the optimal singular control law, which is a…

最优化与控制 · 数学 2026-05-14 Zongxia Liang , Xiaodong Luo , Xiang Yu

In this work, we investigate a security game between an attacker and a defender, originally proposed in \cite{emadi2019security}. As is well known, the combinatorial nature of security games leads to a large cost matrix. Therefore,…

计算机科学与博弈论 · 计算机科学 2020-07-30 HAmid Emadi , Sourabh Bhattacharya

Q-learning is a stochastic approximation version of the classic value iteration. The literature has established that Q-learning suffers from both maximization bias and slower convergence. Recently, multi-step algorithms have shown practical…

机器学习 · 计算机科学 2024-07-03 Antony Vijesh , Shreyas S R

We introduce a novel class of Nash equilibrium seeking dynamics for non-cooperative games with a finite number of players, where the convergence to the Nash equilibrium is bounded by a KL function with a settling time that can be upper…

最优化与控制 · 数学 2020-12-25 Jorge I. Poveda , Miroslav Krstic , Tamer Basar

We introduce a reinforcement learning algorithm designed to identify the fixed points of a given quantum operation. The method iteratively constructs the unitary transformation that maps the computational basis onto the basis of fixed…

量子物理 · 物理学 2025-11-25 María Laura Olivera-Atencio , Jesús Casado-Pascual , Denis Lacroix

We consider zero-sum stochastic games for continuous time Markov decision processes with risk-sensitive average cost criterion. Here the transition and cost rates may be unbounded. We prove the existence of the value of the game and a…

最优化与控制 · 数学 2021-09-21 Mrinal K. Ghosh , Subrata Golui , Chandan Pal , Somnath Pradhan

Quantum Tiq-Taq-Toe is a well-known benchmark and playground for both quantum computing and machine learning. Despite its popularity, no reinforcement learning (RL) methods have been applied to Quantum Tiq-Taq-Toe. Although there has been…

人工智能 · 计算机科学 2024-11-12 Catalin-Viorel Dinu , Thomas Moerland

In this paper, we explore the susceptibility of the independent Q-learning algorithms (a classical and widely used multi-agent reinforcement learning method) to strategic manipulation of sophisticated opponents in normal-form games played…

计算机科学与博弈论 · 计算机科学 2024-07-17 Yuksel Arslantas , Ege Yuceel , Muhammed O. Sayin

Growing advancements in reinforcement learning has led to advancements in control theory. Reinforcement learning has effectively solved the inverted pendulum problem and more recently the double inverted pendulum problem. In reinforcement…

机器学习 · 计算机科学 2021-05-26 Amartya Mukherjee