中文
相关论文

相关论文: MDPs with Energy-Parity Objectives

200 篇论文

We present a memory-bounded optimization approach for solving infinite-horizon decentralized POMDPs. Policies for each agent are represented by stochastic finite state controllers. We formulate the problem of optimizing these policies as a…

人工智能 · 计算机科学 2012-06-26 Christopher Amato , Daniel S Bernstein , Shlomo Zilberstein

Robust Markov decision processes (RMDPs) extend standard Markov decision processes (MDPs) to account for uncertainty in the transition probabilities. RMDPs have an uncertainty set that defines a set of possible transition functions, each of…

计算机科学中的逻辑 · 计算机科学 2026-04-30 Marnix Suilen , Guillermo A. Pérez

Energy Markov Decision Processes (EMDPs) are finite-state Markov decision processes where each transition is assigned an integer counter update and a rational payoff. An EMDP configuration is a pair s(n), where s is a control state and n is…

计算机科学中的逻辑 · 计算机科学 2016-07-05 Tomáš Brázdil , Antonín Kučera , Petr Novotný

Real-world sequential decision making problems commonly involve partial observability, which requires the agent to maintain a memory of history in order to infer the latent states, plan and make good decisions. Coping with partial…

机器学习 · 计算机科学 2022-02-09 Yonathan Efroni , Chi Jin , Akshay Krishnamurthy , Sobhan Miryoosefi

Mean-payoff games (MPGs) are infinite duration two-player zero-sum games played on weighted graphs. Under the hypothesis of perfect information, they admit memoryless optimal strategies for both players and can be solved in…

计算机科学中的逻辑 · 计算机科学 2015-04-14 Paul Hunter , Guillermo A. Pérez , Jean-François Raskin

In this paper, we consider the problem of optimal demand response and energy storage management for a power consuming entity. The entity's objective is to find an optimal control policy for deciding how much load to consume, how much power…

最优化与控制 · 数学 2012-05-22 Longbo Huang , Jean Walrand , Kannan Ramchandran

Battery-less Internet of Things (IoT) devices rely on ambient energy harvesting and therefore require scheduling policies that jointly account for energy intermittency and hard timing constraints. This challenge is especially acute in…

系统与控制 · 电气工程与系统科学 2026-05-20 Shahab Jahanbazi , Mateen Ashraf , Onel L. A. López

We introduce the \emph{submodular objectives chasing problem}, which generalizes many natural and previously-studied problems: a sequence of constrained submodular maximization problems is revealed over time, with both the objective and…

数据结构与算法 · 计算机科学 2025-11-18 Niv Buchbinder , Joseph , Naor , David Wajc

We propose a framework for transferring any existing policy from a potentially unknown source MDP to a target MDP. This framework (1) enables reuse in the target domain of any form of source policy, including classical controllers,…

机器学习 · 计算机科学 2021-01-01 Daniel Graves , Jun Jin , Jun Luo

In this paper, we study turn-based quantitative multiplayer non zero-sum games played on finite graphs with both reachability and safety objectives. In this framework a player with a reachability objective aims at reaching his own goal as…

计算机科学与博弈论 · 计算机科学 2012-05-23 Thomas Brihaye , Véronique Bruyère , Julie De Pril

A challenging category of robotics problems arises when sensing incurs substantial costs. This paper examines settings in which a robot wishes to limit its observations of state, for instance, motivated by specific considerations of energy…

机器人学 · 计算机科学 2023-09-26 Patrick Zhong , Federico Rossi , Dylan A. Shell

We consider synthesis of control policies that maximize the probability of satisfying given temporal logic specifications in unknown, stochastic environments. We model the interaction between the system and its environment as a Markov…

系统与控制 · 计算机科学 2014-05-01 Jie Fu , Ufuk Topcu

A tenet of reinforcement learning is that the agent always observes rewards. However, this is not true in many realistic settings, e.g., a human observer may not always be available to provide rewards, sensors may be limited or…

机器学习 · 计算机科学 2026-03-24 Alireza Kazemipour , Simone Parisi , Matthew E. Taylor , Michael Bowling

We motivate Energy-Based Models (EBMs) as a promising model class for continual learning problems. Instead of tackling continual learning via the use of external memory, growing models, or regularization, EBMs change the underlying training…

机器学习 · 计算机科学 2025-03-05 Shuang Li , Yilun Du , Gido M. van de Ven , Igor Mordatch

It is well-known that the winning region of a parity game with $n$ nodes and $k$ priorities can be computed as a $k$-nested fixpoint of a suitable function; straightforward computation of this nested fixpoint requires…

计算复杂性 · 计算机科学 2021-03-23 Daniel Hausmann , Lutz Schröder

In this paper, for Markov decision processes (MDPs) with unbounded state spaces we present refined upper bounds presented in [Kara et. al. JMLR'23] on finite model approximation errors via optimizing the quantizers used for finite model…

最优化与控制 · 数学 2025-10-15 Osman Bicer , Ali D. Kara , Serdar Yuksel

We study the problem of synthesizing a controller that maximizes the entropy of a partially observable Markov decision process (POMDP) subject to a constraint on the expected total reward. Such a controller minimizes the predictability of a…

最优化与控制 · 数学 2019-09-16 Michael Hibbard , Yagiz Savas , Bo Wu , Takashi Tanaka , Ufuk Topcu

The rapid uptake of renewable energy sources in the electricity grid leads to a demand in load shaping and flexibility. Energy storage devices such as batteries are a key element to provide solutions to these tasks. However, typically a…

最优化与控制 · 数学 2019-07-16 Philipp Sauerteig , Karl Worthmann

Preference-conditioned multi-objective reinforcement learning aims to learn a single policy that captures trade-offs across preferences, but under nonlinear scalarization the uniqueness and continuity of the preference-to-solution…

机器学习 · 计算机科学 2026-05-12 Akihiro Kubo , Kosuke Nakanishi , Shin Ishii

We consider a Reinforcement Learning setup where an agent interacts with an environment in observation-reward-action cycles without any (esp.\ MDP) assumptions on the environment. State aggregation and more generally feature reinforcement…

人工智能 · 计算机科学 2014-07-15 Marcus Hutter