相关论文: On differentiability of reward functionals corresp…
In the paper we study continuous time controlled Markov processes using discrete time controlled Markov processes. We consider long run functionals: average reward per unit time or long run risk sensitive functional. We also investigate…
We consider the expressivity of Markov rewards in sequential decision making under uncertainty. We view reward functions in Markov Decision Processes (MDPs) as a means to characterize desired behaviors of agents. Assuming desired behaviors…
In Markov Decision Processes (MDPs), the reward obtained in a state depends on the properties of the last state and action. This state dependency makes it difficult to reward more interesting long-term behaviors, such as always closing a…
Reward is the driving force for reinforcement-learning agents. This paper is dedicated to understanding the expressivity of reward as a way to capture tasks that we would want an agent to perform. We frame this study around three new…
We consider a Markov control model in discrete time with countable both state space and action space. Using the value function of a suitable long-run average reward problem, we study various reachability/controllability problems. First, we…
This paper studies the expected value of multiplicative rewards, where rewards obtained in each step are multiplied (instead of the usual addition), in Markov chains (MCs) and Markov decision processes (MDPs). One of the key differences to…
Controlled discrete time Markov processes are studied first with long run general discounting functional. It is shown that optimal strategies for average reward per unit time problem are also optimal for average generally discounting…
Markov decision processes are typically used for sequential decision making under uncertainty. For many aspects however, ranging from constrained or safe specifications to various kinds of temporal (non-Markovian) dependencies in task and…
Markov decision processes (MDPs) are standard models for probabilistic systems with non-deterministic behaviours. Long-run average rewards provide a mathematically elegant formalism for expressing long term performance. Value iteration (VI)…
In the paper average reward per unit time and average risk sensitive reward functionals are considered for controlled nonhomogeneous Markov processes. Existence of solutions to suitable Bellman equations is shown. Continuity of the value…
A Markov decision process can be parameterized by a transition kernel and a reward function. Both play essential roles in the study of reinforcement learning as evidenced by their presence in the Bellman equations. In our inquiry of various…
In the paper we study dependence of long run functionals and limit characteristics assuming that Borel measurable Markov controls converge pointwise. We consider two kinds of functionals: average cost per unit time and long run risk…
The problem of p-th moment stability for time-varying stochastic time-delay systems with Markovian switching is investigated in this paper. Some novel stability criteria are obtained by applying the generalized Razumikhin and Krasovskii…
The problem of reward design examines the interaction between a leader and a follower, where the leader aims to shape the follower's behavior to maximize the leader's payoff by modifying the follower's reward function. Current approaches to…
We explore properties of the value function and existence of optimal stopping times for functionals with discontinuities related to the boundary of an open (possibly unbounded) set $\mathcal{O}$. The stopping horizon is either random, equal…
In this paper we introduce and solve a class of optimal stopping problems of recursive type. In particular, the stopping payoff depends directly on the value function of the problem itself. In a multi-dimensional Markovian setting we show…
We consider a renewal-reward process with multivariate rewards. Such a process is constructed from an i.i.d.\ sequence of time periods, to each of which there is associated a multivariate reward vector. The rewards in each time period may…
Many Reinforcement Learning algorithms assume a Markov reward function to guarantee optimality. However, not all reward functions are Markov. This paper proposes a framework for mapping non-Markov reward functions into equivalent Markov…
Inverse reinforcement learning attempts to reconstruct the reward function in a Markov decision problem, using observations of agent actions. As already observed in Russell [1998] the problem is ill-posed, and the reward function is not…
Stochastically monotone Markov chains arise in many applied domains, especially in the setting of queues and storage systems. Poisson's equation is a key tool for analyzing additive functionals of such models, such as cumulative sums of…