中文
相关论文

相关论文: HSVI-based Online Minimax Strategies for Partially…

200 篇论文

Stochastic games are a well established model for multi-agent sequential decision making under uncertainty. In practical applications, though, agents often have only partial observability of their environment. Furthermore, agents…

计算机科学与博弈论 · 计算机科学 2024-07-02 Rui Yan , Gabriel Santos , Gethin Norman , David Parker , Marta Kwiatkowska

State-of-the-art methods for solving 2-player zero-sum imperfect information games rely on linear programming or regret minimization, though not on dynamic programming (DP) or heuristic search (HS), while the latter are often at the core of…

人工智能 · 计算机科学 2022-10-27 Aurélien Delage , Olivier Buffet , Jilles S. Dibangoye , Abdallah Saffidine

Game-theoretic agents must make plans that optimally gather information about their opponents. These problems are modeled by partially observable stochastic games (POSGs), but planning in fully continuous POSGs is intractable without heavy…

计算机科学与博弈论 · 计算机科学 2025-06-03 Mel Krusniak , Hang Xu , Parker Palermo , Forrest Laine

We study episodic two-player zero-sum Markov games (MGs) in the offline setting, where the goal is to find an approximate Nash equilibrium (NE) policy pair based on a dataset collected a priori. When the dataset does not have uniform…

机器学习 · 计算机科学 2023-01-02 Han Zhong , Wei Xiong , Jiyuan Tan , Liwei Wang , Tong Zhang , Zhaoran Wang , Zhuoran Yang

In two-player finite-state stochastic games of partial observation on graphs, in every state of the graph, the players simultaneously choose an action, and their joint actions determine a probability distribution over the successor states.…

计算机科学与博弈论 · 计算机科学 2011-07-13 Krishnendu Chatterjee , Laurent Doyen

Partially observable stochastic games provide a rich mathematical paradigm for modeling multi-agent dynamic decision making under uncertainty and partial information. However, they generally do not admit closed-form solutions and are…

最优化与控制 · 数学 2020-04-15 Yanling Chang , Chelsea C. White

Dynamic programming and heuristic search are at the core of state-of-the-art solvers for sequential decision-making problems. In partially observable or collaborative settings (\eg, POMDPs and Dec-POMDPs), this requires introducing an…

计算机科学与博弈论 · 计算机科学 2022-11-16 Aurélien Delage , Olivier Buffet , Jilles Dibangoye

Technology development efforts in autonomy and cyber-defense have been evolving independently of each other, over the past decade. In this paper, we report our ongoing effort to integrate these two presently distinct areas into a single…

计算机科学与博弈论 · 计算机科学 2020-02-07 Mohamadreza Ahmadi , Arun A. Viswanathan , Michel D. Ingham , Kymie Tan , Aaron D. Ames

The goal of agents in multi-agent environments is to maximize total reward against the opposing agents that are encountered. Following a game-theoretic solution concept, such as Nash equilibrium, may obtain a strong performance in some…

计算机科学与博弈论 · 计算机科学 2026-01-05 Sam Ganzfried

We consider a scenario where a team of two unmanned aerial vehicles (UAVs) pursue an evader UAV within an urban environment. Each agent has a limited view of their environment where buildings can occlude their field-of-view. Additionally,…

多智能体系统 · 计算机科学 2025-11-13 Addison Kalanther , Daniel Bostwick , Chinmay Maheshwari , Shankar Sastry

Robots deployed to the real world must be able to interact with other agents in their environment. Dynamic game theory provides a powerful mathematical framework for modeling scenarios in which agents have individual objectives and…

This letter studies multi-agent reinforcement learning in partially observable Markov potential games. Solving this problem is challenging due to partial observability, decentralized information, and the curse of dimensionality. First, to…

多智能体系统 · 计算机科学 2026-04-02 Wonseok Yang , Thinh T. Doan

Adversarial decision-making in partially observable multi-agent systems requires sophisticated strategies for both deception and counter-deception. This paper presents a sequential hypothesis testing (SHT)-driven framework that captures the…

最优化与控制 · 数学 2026-04-14 Haosheng Zhou , Daniel Ralston , Xu Yang , Ruimeng Hu

In imperfect-information games, agents must make decisions based on partial knowledge of the game state. The Belief Stochastic Game model addresses this challenge by delegating state estimation to the game model itself. This allows agents…

人工智能 · 计算机科学 2025-08-20 Achille Morenville , Éric Piette

In this article we analyze a partial-information Nash Q-learning algorithm for a general 2-player stochastic game. Partial information refers to the setting where a player does not know the strategy or the actions taken by the opposing…

计算机科学与博弈论 · 计算机科学 2023-02-22 Negash Medhin , Andrew Papanicolaou , Marwen Zrida

We study a pursuit-evasion game between two players with car-like dynamics and sensing limitations by formalizing it as a partially observable stochastic zero-sum game. The partial observability caused by the sensing constraints is…

机器人学 · 计算机科学 2025-06-17 Burak M. Gonultas , Volkan Isler

A multi-agent system operates in an uncertain environment about which agents have different and time varying beliefs that, as time progresses, converge to a common belief. A global utility function that depends on the realized state of the…

计算机科学与博弈论 · 计算机科学 2016-02-08 Ceyhun Eksin , Alejandro Ribeiro

We present a framework that incorporates the idea of bounded rationality into dynamic stochastic pursuit-evasion games. The solution of a stochastic game is characterized, in general, by its (Nash) equilibria in feedback form. However,…

系统与控制 · 电气工程与系统科学 2020-03-17 Yue Guan , Dipankar Maity , Christopher M. Kroninger , Panagiotis Tsiotras

Many real-world decision problems involve the interaction of multiple self-interested agents with limited sensing ability. The partially observable stochastic game (POSG) provides a mathematical framework for modeling these problems,…

计算机科学与博弈论 · 计算机科学 2024-10-30 Tyler Becker , Zachary Sunberg

We address the online linear optimization problem when the actions of the forecaster are represented by binary vectors. Our goal is to understand the magnitude of the minimax regret for the worst possible set of actions. We study the…

机器学习 · 统计学 2011-05-25 Jean-Yves Audibert , Sebastien Bubeck , Gabor Lugosi
‹ 上一页 1 2 3 10 下一页 ›