中文
相关论文

相关论文: HSVI can solve zero-sum Partially Observable Stoch…

200 篇论文

Dynamic programming and heuristic search are at the core of state-of-the-art solvers for sequential decision-making problems. In partially observable or collaborative settings (\eg, POMDPs and Dec-POMDPs), this requires introducing an…

计算机科学与博弈论 · 计算机科学 2022-11-16 Aurélien Delage , Olivier Buffet , Jilles Dibangoye

We present a novel framework for {\epsilon}-optimally solving two-player zero-sum partially observable stochastic games (zs-POSGs). These games pose a major challenge due to the absence of a principled connection with dynamic programming…

计算机科学与博弈论 · 计算机科学 2025-11-17 Erwan Christian Escudie , Matthia Sabatelli , Olivier Buffet , Jilles Steeve Dibangoye

Many security and other real-world situations are dynamic in nature and can be modelled as strictly competitive (or zero-sum) dynamic games. In these domains, agents perform actions to affect the environment and receive observations --…

计算机科学与博弈论 · 计算机科学 2020-10-23 Karel Horák , Branislav Bošanský , Vojtěch Kovařík , Christopher Kiekintveld

Many non-trivial sequential decision-making problems are efficiently solved by relying on Bellman's optimality principle, i.e., exploiting the fact that sub-problems are nested recursively within the original problem. Here we show how it…

人工智能 · 计算机科学 2022-11-16 Olivier Buffet , Jilles Dibangoye , Aurélien Delage , Abdallah Saffidine , Vincent Thomas

We present a novel POMDP planning algorithm called heuristic search value iteration (HSVI).HSVI is an anytime algorithm that returns a policy and a provable bound on its regret with respect to the optimal policy. HSVI gets its power by…

人工智能 · 计算机科学 2012-07-19 Trey Smith , Reid Simmons

Many real-world decision problems involve the interaction of multiple self-interested agents with limited sensing ability. The partially observable stochastic game (POSG) provides a mathematical framework for modeling these problems,…

计算机科学与博弈论 · 计算机科学 2024-10-30 Tyler Becker , Zachary Sunberg

We consider a variant of continuous-state partially-observable stochastic games with neural perception mechanisms and an asymmetric information structure. One agent has partial information, with the observation function implemented as a…

计算机科学与博弈论 · 计算机科学 2024-04-17 Rui Yan , Gabriel Santos , Gethin Norman , David Parker , Marta Kwiatkowska

A recent method for solving zero-sum partially observable stochastic games (zs-POSGs) embeds the original game into a new one called the occupancy Markov game. This reformulation allows applying Bellman's principle of optimality to solve…

计算机科学与博弈论 · 计算机科学 2024-06-04 Erwan Escudie , Matthia Sabatelli , Jilles Dibangoye

Stochastic games are a well established model for multi-agent sequential decision making under uncertainty. In practical applications, though, agents often have only partial observability of their environment. Furthermore, agents…

计算机科学与博弈论 · 计算机科学 2024-07-02 Rui Yan , Gabriel Santos , Gethin Norman , David Parker , Marta Kwiatkowska

This paper considers the problem of two-player zero-sum stochastic differential game with both players adopting impulse controls in finite horizon under rather weak assumptions on the cost functions ($c$ and $\chi$ not decreasing in time).…

最优化与控制 · 数学 2018-09-26 Brahim El Asri , Sehail Mazid

Zero-sum stochastic games provide a rich model for competitive decision making. However, under general forms of state uncertainty as considered in the Partially Observable Stochastic Game (POSG), such decision making problems are still not…

人工智能 · 计算机科学 2016-06-23 Auke J. Wiggers , Frans A. Oliehoek , Diederik M. Roijers

Unlike Poker where the action space $\mathcal{A}$ is discrete, differential games in the physical world often have continuous action spaces not amenable to discrete abstraction, rendering no-regret algorithms with…

计算机科学与博弈论 · 计算机科学 2025-02-17 Mukesh Ghimire , Zhe Xu , Yi Ren

A general model for zero-sum stochastic games with asymmetric information is considered. In this model, each player's information at each time can be divided into a common information part and a private information part. Under certain…

系统与控制 · 电气工程与系统科学 2019-12-25 Dhruva Kartik , Ashutosh Nayyar

In the present paper, we study a two-player zero-sum deterministic differential game with both players adopting impulse controls, in infinite time horizon, under rather weak assumptions on the cost functions. We prove by means of the…

最优化与控制 · 数学 2021-01-29 Brahim El Asri , Hafid Lalioui , Sehail Mazid

Information theory has been very successful in obtaining performance limits for various problems such as communication, compression and hypothesis testing. Likewise, stochastic control theory provides a characterization of optimal policies…

信息论 · 计算机科学 2018-10-15 Dhruva Kartik , Ekraam Sabir , Urbashi Mitra , Prem Natarajan

Game-theoretic agents must make plans that optimally gather information about their opponents. These problems are modeled by partially observable stochastic games (POSGs), but planning in fully continuous POSGs is intractable without heavy…

计算机科学与博弈论 · 计算机科学 2025-06-03 Mel Krusniak , Hang Xu , Parker Palermo , Forrest Laine

Pursuit-evasion scenarios appear widely in robotics, security domains, and many other real-world situations. We focus on two-player pursuit-evasion games with concurrent moves, infinite horizon, and discounted rewards. We assume that the…

计算机科学与博弈论 · 计算机科学 2016-08-05 Karel Horák , Branislav Bošanský

Stochastic games are an important class of problems that generalize Markov decision processes to game theoretic scenarios. We consider finite state two-player zero-sum stochastic games over an infinite time horizon with discounted rewards.…

最优化与控制 · 数学 2008-06-17 Parikshit Shah , Pablo A. Parrilo

We consider a two-player zero-sum stochastic differential game in which one of the players has a private information on the game. Both players observe each other, so that the non-informed player can try to guess his missing information. Our…

概率论 · 数学 2011-06-15 Christine Grün

Multi-agent planning and reinforcement learning can be challenging when agents cannot see the state of the world or communicate with each other due to communication costs, latency, or noise. Partially Observable Stochastic Games (POSGs)…

多智能体系统 · 计算机科学 2024-12-20 Rafael F. Cunha , Jacopo Castellini , Johan Peralez , Jilles S. Dibangoye
‹ 上一页 1 2 3 10 下一页 ›