Related papers: On Occupation Time for On-Off Processes with Multi…
We study the problem of off-policy policy optimization in Markov decision processes, and develop a novel off-policy policy gradient method. Prior off-policy policy gradient approaches have generally ignored the mismatch between the…
Markov switching models are a popular family of models that introduces time-variation in the parameters in the form of their state- or regime-specific values. Importantly, this time-variation is governed by a discrete-valued latent…
In this paper we study coupled fully non-local equations, where a linear non-local operator jointly acts on the time and space variables. We establish existence and uniqueness of the solution. A maximum principle is proved and used to…
This paper introduces an analytical formula for the fractional-order conditional moments of nonlinear drift constant elasticity of variance (NLD-CEV) processes under regime switching, governed by continuous-time finite-state irreducible…
Policy gradient methods are widely adopted reinforcement learning algorithms for tasks with continuous action spaces. These methods succeeded in many application domains, however, because of their notorious sample inefficiency their use…
Several Markovian process calculi have been proposed in the literature, which differ from each other for various aspects. With regard to the action representation, we distinguish between integrated-time Markovian process calculi, in which…
This paper is devoted to the study of a stochastic process obtained by random switching between a finite collection of vector fields. Such processes have recently been the focus of much attention in the case where the switching times are…
Let $(X_t)_{t \geq 0}$ be a continuous time Markov process on some metric space $M,$ leaving invariant a closed subset $M_0 \subset M,$ called the {\em extinction set}. We give general conditions ensuring either "Stochastic persistence"…
Model-based methods have recently shown great potential for off-policy evaluation (OPE); offline trajectories induced by behavioral policies are fitted to transitions of Markov decision processes (MDPs), which are used to rollout simulated…
This paper studies a class of optimal multiple stopping problems driven by L\'evy processes. Our model allows for a negative effective discount rate, which arises in a number of financial applications, including stock loans and real…
Reinforcement learning (RL) tasks are typically framed as Markov Decision Processes (MDPs), assuming that decisions are made at fixed time intervals. However, many applications of great importance, including healthcare, do not satisfy this…
We consider a queuing model with the workload evolving between consecutive i.i.d. exponential timers $\{e_q^{(i)}\}_{i=1,2,...}$ according to a spectrally positive L\'{e}vy process $Y(t)$ which is reflected at 0. When the exponential clock…
In this paper we present elementary computations for some Markov modulated counting processes, also called counting processes with regime switching. Regime switching has become an increasingly popular concept in many branches of science. In…
We consider a finite state discrete time process X. Without loss of generality the finite state space can be identified with the set of unit vectors {e1, e2, . . . , eN} with ei = (0, . . . , 0, 1, 0, . . . , 0)0 2 RN. For a Markov chain…
This work studies the problem of batch off-policy evaluation for Reinforcement Learning in partially observable environments. Off-policy evaluation under partial observability is inherently prone to bias, with risk of arbitrarily large…
This paper considers optimization over multiple renewal systems coupled by time average constraints. These systems act asynchronously over variable length frames. For each system, at the beginning of each renewal frame, it chooses an action…
This paper studies the problem of optimally extracting nonrenewable natural resources. Taking into account the fact that the market values of the main natural resources i.e. oil, natural gas, copper,..., etc, fluctuate randomly following…
The distribution of the "mixing time" or the "time to stationarity" in a discrete time irreducible Markov chain, starting in state i, can be defined as the number of trials to reach a state sampled from the stationary distribution of the…
The main topic of these notes are Markov loops, studied in the context of continuous time Markov chains on discrete state spaces. We refer to [1] and [2] for the short "history" of the subject. In contrast with these references, symmetry is…
Batch offline data have been shown considerably beneficial for reinforcement learning. Their benefit is further amplified by upsampling with generative models. In this paper, we consider a novel opportunity where interaction with…