中文
相关论文

相关论文: Exploring an Infinite Space with Finite Memory Sco…

200 篇论文

We study 2-player zero-sum concurrent (i.e., simultaneous move) stochastic B\"uchi games and Transience games on countable graphs. Two players, Max and Min, seek respectively to maximize and minimize the probability of satisfying the game…

计算机科学与博弈论 · 计算机科学 2025-06-23 Stefan Kiefer , Richard Mayr , Mahsa Shirmohammadi , Patrick Totzke

We consider the following distributed pursuit-evasion problem. A team of mobile agents called searchers starts at an arbitrary node of an unknown $n$-node network. Their goal is to execute a search strategy that guarantees capturing a fast…

离散数学 · 计算机科学 2021-01-19 Dariusz Dereniowski , Dorota Urbańska

This paper studies finite-time safety and reach-avoid verification for stochastic discrete-time dynamical systems. The aim is to ascertain lower and upper bounds of the probability that, within a predefined finite-time horizon, a system…

系统与控制 · 电气工程与系统科学 2025-10-22 Bai Xue

We study variants of regular infinite games where the strict alternation of moves between the two players is subject to modifications. The second player may postpone a move for a finite number of steps, or, in other words, exploit in his…

形式语言与自动机理论 · 计算机科学 2015-07-01 Michael Holtmann , Lukasz Kaiser , Wolfgang Thomas

We study an identification problem in multi-armed bandits. In each round a learner selects one of $K$ arms and observes its reward, with the goal of eventually identifying an arm that will perform best at a {\it future} time. In adversarial…

机器学习 · 计算机科学 2026-03-03 Nataly Brukhim , Nicolò Cesa-Bianchi , Carlo Ciliberto

This paper aims to put forward the concept that learning to take safe actions in unknown environments, even with probability one guarantees, can be achieved without the need for an unbounded number of exploratory trials, provided that one…

机器学习 · 计算机科学 2021-04-01 Agustin Castellano , Juan Bazerque , Enrique Mallada

We revisit the problem of searching for a target at an unknown location on a line when given upper and lower bounds on the distance D that separates the initial position of the searcher from the target. Prior to this work, only asymptotic…

数据结构与算法 · 计算机科学 2013-10-04 Prosenjit Bose , Jean-Lou De Carufel , Stephane Durocher

We consider a team of {\em autonomous weak robots} that are endowed with visibility sensors and motion actuators. Autonomous means that the team cannot rely on any kind of central coordination mechanism or scheduler. By weak we mean that…

分布式、并行与集群计算 · 计算机科学 2012-03-08 Stéphane Devismes , Anissa Lamani , Franck Petit , Pascal Raymond , Sébastien Tixeuil

We study the repeated principal-agent bandit game, where the principal indirectly interacts with the unknown environment by proposing incentives for the agent to play arms. Most existing work assumes the agent has full knowledge of the…

机器学习 · 计算机科学 2025-06-03 Junyan Liu , Lillian J. Ratliff

We establish the existence of optimal scheduling strategies for time-bounded reachability in continuous-time Markov decision processes, and of co-optimal strategies for continuous-time Markov games. Furthermore, we show that optimal control…

形式语言与自动机理论 · 计算机科学 2010-06-07 Markus Rabe , Sven Schewe

In this paper we consider stochastic multiarmed bandit problems. Recently a policy, DMED, is proposed and proved to achieve the asymptotic bound for the model that each reward distribution is supported in a known bounded interval, e.g.…

统计理论 · 数学 2012-02-20 Junya Honda , Akimichi Takemura

This paper considers a reach-avoid differential game in three-dimensional space with four equal-speed players. A plane divides the game space into a play subspace and a goal subspace. The evader aims at entering the goal subspace while…

最优化与控制 · 数学 2019-04-08 Rui Yan , Zongying Shi , Yisheng Zhong

In this work we create agents that can perform well beyond a single, individual task, that exhibit much wider generalisation of behaviour to a massive, rich space of challenges. We define a universe of tasks within an environment domain and…

Multi-dimensional mean-payoff and energy games provide the mathematical foundation for the quantitative study of reactive systems, and play a central role in the emerging quantitative theory of verification and synthesis. In this work, we…

计算机科学与博弈论 · 计算机科学 2014-11-04 Krishnendu Chatterjee , Mickael Randour , Jean-François Raskin

Two identical anonymous mobile agents have to meet at a node of the infinite oriented grid whose nodes are unlabeled. This problem is known as rendezvous. The agents execute the same deterministic algorithm. Time is divided into rounds, and…

分布式、并行与集群计算 · 计算机科学 2024-07-23 Younan Gao , Andrzej Pelc

In this paper we consider the problem of learning the optimal policy for uncontrolled restless bandit problems. In an uncontrolled restless bandit problem, there is a finite set of arms, each of which when pulled yields a positive reward.…

最优化与控制 · 数学 2015-01-30 Cem Tekin , Mingyan Liu

In this paper, we consider the distributed stochastic multi-armed bandit problem, where a global arm set can be accessed by multiple players independently. The players are allowed to exchange their history of observations with each other at…

机器学习 · 计算机科学 2020-02-13 Shuang Liu , Cheng Chen , Zhihua Zhang

Priced timed games are optimal-cost reachability games played between two players---the controller and the environment---by moving a token along the edges of infinite graphs of configurations of priced timed automata. The goal of the…

计算机科学中的逻辑 · 计算机科学 2015-07-22 Shibashis Guha , Shankara Narayanan Krishna , Lakshmi Manasa , Ashutosh Trivedi

We consider finite-state Markov decision processes with the combined Energy-MeanPayoff objective. The controller tries to avoid running out of energy while simultaneously attaining a strictly positive mean payoff in a second dimension. We…

计算机科学与博弈论 · 计算机科学 2025-10-13 Mohan Dantam , Richard Mayr

In this paper, we consider a bandit problem in which there are a number of groups each consisting of infinitely many arms. Whenever a new arm is requested from a given group, its mean reward is drawn from an unknown reservoir distribution…

机器学习 · 统计学 2023-02-02 Ivan Lau , Yan Hao Ling , Mayank Shrivastava , Jonathan Scarlett