中文
相关论文

相关论文: On differentiability of reward functionals corresp…

200 篇论文

In the paper we study continuous time controlled Markov processes using discrete time controlled Markov processes. We consider long run functionals: average reward per unit time or long run risk sensitive functional. We also investigate…

最优化与控制 · 数学 2025-08-12 Lukasz Stettner

We consider the expressivity of Markov rewards in sequential decision making under uncertainty. We view reward functions in Markov Decision Processes (MDPs) as a means to characterize desired behaviors of agents. Assuming desired behaviors…

人工智能 · 计算机科学 2023-07-25 Shuwa Miura

In Markov Decision Processes (MDPs), the reward obtained in a state depends on the properties of the last state and action. This state dependency makes it difficult to reward more interesting long-term behaviors, such as always closing a…

人工智能 · 计算机科学 2017-06-27 Ronen Brafman , Giuseppe De Giacomo , Fabio Patrizi

Reward is the driving force for reinforcement-learning agents. This paper is dedicated to understanding the expressivity of reward as a way to capture tasks that we would want an agent to perform. We frame this study around three new…

机器学习 · 计算机科学 2022-01-19 David Abel , Will Dabney , Anna Harutyunyan , Mark K. Ho , Michael L. Littman , Doina Precup , Satinder Singh

We consider a Markov control model in discrete time with countable both state space and action space. Using the value function of a suitable long-run average reward problem, we study various reachability/controllability problems. First, we…

最优化与控制 · 数学 2024-06-05 Daniel Avila , Mauricio Junca

This paper studies the expected value of multiplicative rewards, where rewards obtained in each step are multiplied (instead of the usual addition), in Markov chains (MCs) and Markov decision processes (MDPs). One of the key differences to…

计算机科学中的逻辑 · 计算机科学 2025-06-24 Christel Baier , Krishnendu Chatterjee , Tobias Meggendorfer , Jakob Piribauer

Controlled discrete time Markov processes are studied first with long run general discounting functional. It is shown that optimal strategies for average reward per unit time problem are also optimal for average generally discounting…

最优化与控制 · 数学 2023-06-27 Łukasz Stettner

Markov decision processes are typically used for sequential decision making under uncertainty. For many aspects however, ranging from constrained or safe specifications to various kinds of temporal (non-Markovian) dependencies in task and…

人工智能 · 计算机科学 2021-11-10 Nicky Lenaers , Martijn van Otterlo

Markov decision processes (MDPs) are standard models for probabilistic systems with non-deterministic behaviours. Long-run average rewards provide a mathematically elegant formalism for expressing long term performance. Value iteration (VI)…

系统与控制 · 计算机科学 2017-09-01 Pranav Ashok , Krishnendu Chatterjee , Przemyslaw Daca , Jan Křetínský , Tobias Meggendorfer

In the paper average reward per unit time and average risk sensitive reward functionals are considered for controlled nonhomogeneous Markov processes. Existence of solutions to suitable Bellman equations is shown. Continuity of the value…

最优化与控制 · 数学 2025-06-19 Łukasz Stettner

A Markov decision process can be parameterized by a transition kernel and a reward function. Both play essential roles in the study of reinforcement learning as evidenced by their presence in the Bellman equations. In our inquiry of various…

机器学习 · 计算机科学 2023-09-04 Falcon Z. Dai

In the paper we study dependence of long run functionals and limit characteristics assuming that Borel measurable Markov controls converge pointwise. We consider two kinds of functionals: average cost per unit time and long run risk…

概率论 · 数学 2024-12-03 Lukasz Stettner

The problem of p-th moment stability for time-varying stochastic time-delay systems with Markovian switching is investigated in this paper. Some novel stability criteria are obtained by applying the generalized Razumikhin and Krasovskii…

动力系统 · 数学 2016-07-11 Bin Zhou , Weiwei Luo

The problem of reward design examines the interaction between a leader and a follower, where the leader aims to shape the follower's behavior to maximize the leader's payoff by modifying the follower's reward function. Current approaches to…

最优化与控制 · 数学 2024-06-10 Shuo Wu , Haoxiang Ma , Jie Fu , Shuo Han

We explore properties of the value function and existence of optimal stopping times for functionals with discontinuities related to the boundary of an open (possibly unbounded) set $\mathcal{O}$. The stopping horizon is either random, equal…

最优化与控制 · 数学 2017-01-11 Jan Palczewski , Lukasz Stettner

In this paper we introduce and solve a class of optimal stopping problems of recursive type. In particular, the stopping payoff depends directly on the value function of the problem itself. In a multi-dimensional Markovian setting we show…

最优化与控制 · 数学 2021-06-23 Katia Colaneri , Tiziano De Angelis

We consider a renewal-reward process with multivariate rewards. Such a process is constructed from an i.i.d.\ sequence of time periods, to each of which there is associated a multivariate reward vector. The rewards in each time period may…

概率论 · 数学 2014-08-08 Brendan Patch , Yoni Nazarathy , Thomas Taimre

Many Reinforcement Learning algorithms assume a Markov reward function to guarantee optimality. However, not all reward functions are Markov. This paper proposes a framework for mapping non-Markov reward functions into equivalent Markov…

机器学习 · 计算机科学 2024-08-19 Gregory Hyde , Eugene Santos

Inverse reinforcement learning attempts to reconstruct the reward function in a Markov decision problem, using observations of agent actions. As already observed in Russell [1998] the problem is ill-posed, and the reward function is not…

机器学习 · 计算机科学 2021-11-09 Haoyang Cao , Samuel N. Cohen , Lukasz Szpruch

Stochastically monotone Markov chains arise in many applied domains, especially in the setting of queues and storage systems. Poisson's equation is a key tool for analyzing additive functionals of such models, such as cumulative sums of…

概率论 · 数学 2022-02-23 Peter W. Glynn , Alex Infanger
‹ 上一页 1 2 3 10 下一页 ›