中文
相关论文

相关论文: Min-Max Q-Learning for Multi-Player Pursuit-Evasio…

200 篇论文

Gradient-based learning in multi-agent systems is difficult because the gradient derives from a first-order model which does not account for the interaction between agents' learning processes. LOLA (arXiv:1709.04326) accounts for this by…

机器学习 · 计算机科学 2023-12-12 Tim Cooijmans , Milad Aghajohari , Aaron Courville

The pursuit-evasion game in Smart City brings a profound impact on the Multi-vehicle Pursuit (MVP) problem, when police cars cooperatively pursue suspected vehicles. Existing studies on the MVP problems tend to set evading vehicles to move…

多智能体系统 · 计算机科学 2022-10-25 Qinwen Wang , Xinhang Li , Zheng Yuan , Yiying Yang , Chen Xu , Lin Zhang

We address the problem of agile 1v1 quadrotor pursuit-evasion, where a pursuer and an evader learn to outmaneuver each other through reinforcement learning (RL). Such settings face two major challenges: non-stationarity, since each agent's…

机器人学 · 计算机科学 2025-09-16 Alejandro Sanchez Roncero , Yixi Cai , Olov Andersson , Petter Ogren

We consider the Reinforcement Learning problem of controlling an unknown dynamical system to maximise the long-term average reward along a single trajectory. Most of the literature considers system interactions that occur in discrete time…

人工智能 · 计算机科学 2023-09-07 Lorenzo Croissant , Marc Abeille , Bruno Bouchard

A very successful model for simulating emergency evacuation is the social-force model. At the heart of the model is the self-driven force that is applied to an agent and is directed towards the exit. However, it is not clear if the…

机器学习 · 计算机科学 2021-03-09 Yihao Zhang , Zhaojie Chai , George Lykotrafitis

This paper studies the problem of multi-robot pursuit of how to coordinate a group of defending robots to capture a faster attacker before it enters a protected area. Such operation for defending robots is challenging due to the unknown…

机器人学 · 计算机科学 2024-11-01 Jinyong Chen , Rui Zhou , Zhaozong Wang , Yunjie Zhang , Guibin Sun

We consider a variant of pursuit-evasion games where a single defender is tasked to defend a static target from a sequence of periodically arriving intruders. The intruders' objective is to breach the boundary of a circular target without…

最优化与控制 · 数学 2023-03-13 Arman Pourghorban , Dipankar Maity

We motivate and propose a new model for non-cooperative Markov game which considers the interactions of risk-aware players. This model characterizes the time-consistent dynamic "risk" from both stochastic state transitions (inherent to the…

计算机科学与博弈论 · 计算机科学 2019-11-22 Wenjie Huang , Pham Viet Hai , William B. Haskell

Model-free Reinforcement Learning (RL) algorithms such as Q-learning [Watkins, Dayan 92] have been widely used in practice and can achieve human level performance in applications such as video games [Mnih et al. 15]. Recently, equipped with…

机器学习 · 计算机科学 2019-05-03 Zhao Song , Wen Sun

Consider a two-player zero-sum stochastic game where the transition function can be embedded in a given feature space. We propose a two-player Q-learning algorithm for approximating the Nash equilibrium strategy via sampling. The algorithm…

机器学习 · 计算机科学 2019-06-04 Zeyu Jia , Lin F. Yang , Mengdi Wang

In multi-agent settings, game theory is a natural framework for describing the strategic interactions of agents whose objectives depend upon one another's behavior. Trajectory games capture these complex effects by design. In competitive…

计算机科学与博弈论 · 计算机科学 2022-05-04 Lasse Peters , David Fridovich-Keil , Laura Ferranti , Cyrill Stachniss , Javier Alonso-Mora , Forrest Laine

This paper addresses zero-sum ``turn'' games, in which only one player can make decisions at each state. We show that pure saddle-point state-feedback policies for turn games can be constructed from dynamic programming fixed-point equations…

系统与控制 · 电气工程与系统科学 2025-09-18 Sean Anderson , Chris Darken , João Hespanha

Robots playing games that humans are adept in is a challenge. We studied robotic agents playing Chain Catch game as a Multi-Agent System (MAS). Our game starts with a traditional Catch game similar to Pursuit evasion, and further extends it…

多智能体系统 · 计算机科学 2016-02-23 Garima Agrawal , Kamalakar Karlapalem

This paper considers an M-pursuer N-evader scenario involving virtual targets. The virtual targets serve as an intermediary target for the pursuers, allowing the pursuers to delay their final assignment to the evaders. However, upon…

最优化与控制 · 数学 2023-12-21 Isaac E. Weintraub , Alexander Von Moll , David W. Casbeer , Satyanarayana G. Manyam

Applying Q-learning to high-dimensional or continuous action spaces can be difficult due to the required maximization over the set of possible actions. Motivated by techniques from amortized inference, we replace the expensive maximization…

机器学习 · 计算机科学 2020-01-23 Tom Van de Wiele , David Warde-Farley , Andriy Mnih , Volodymyr Mnih

Guided exploration with expert demonstrations improves data efficiency for reinforcement learning, but current algorithms often overuse expert information. We propose a novel algorithm to speed up Q-learning with the help of a limited…

机器学习 · 计算机科学 2022-10-06 Fengdi Che , Xiru Zhu , Doina Precup , David Meger , Gregory Dudek

The design and testing of supervised machine learning models combine two fundamental distributions: (1) the training data distribution (2) the testing data distribution. Although these two distributions are identical and identifiable when…

机器学习 · 计算机科学 2021-03-19 Peyman Tavallali , Hamed Hamze Bajgiran , Danial J. Esaid , Houman Owhadi

The goal of this paper is to propose a new Q-learning algorithm with a dummy adversarial player, which is called dummy adversarial Q-learning (DAQ), that can effectively regulate the overestimation bias in standard Q-learning. With the…

机器学习 · 计算机科学 2024-10-01 HyeAnn Lee , Donghwan Lee

This article investigates the optimal control problem with disturbance rejection for discrete-time multi-agent systems under cooperative and non-cooperative graphical games frameworks. Given the practical challenges of obtaining accurate…

系统与控制 · 电气工程与系统科学 2025-04-11 Xinyang Wang , Martin Guay , Shimin Wang , Hongwei Zhang

This paper studies a novel encirclement guaranteed cooperative pursuit problem involving $N$ pursuers and a single evader in an unbounded two-dimensional game domain. Throughout the game, the pursuers are required to maintain encirclement…

多智能体系统 · 计算机科学 2024-04-08 Chen Wang , Hua Chen , Jia Pan , Wei Zhang