Related papers: Mimicking and Conditional Control with Hard Killin…
To describe and analyze the dynamics of Self-Organized Criticality (SOC) systems, a four-state continuous-time Markov model is proposed in this paper. Different to computer simulation or numeric experimental approaches commonly employed for…
We provide results demonstrating the smoothness of some marginal log-linear parameterizations for distributions on multi-way contingency tables. First we give an analytical relationship between log-linear parameters defined within different…
Many applications -- including power systems, robotics, and economics -- involve a dynamical system interacting with a stochastic and hard-to-model environment. We adopt a reinforcement learning approach to control such systems.…
Many chemical processes exhibit diverse timescale dynamics with a strong coupling between timescale sensitive variables. Model predictive control with a non-uniformly spaced optimisation horizon is an effective approach to multi-timescale…
We describe an abstract control-theoretic framework in which the validity of the dynamic programming principle can be established in continuous time by a verification of a small number of structural properties. As an application we treat…
We exhibit conditions under which the flow of marginal distributions of a discontinuous semimartingale $\xi$ can be matched by a Markov process, whose infinitesimal generator is expressed in terms of the local characteristics of $\xi$. Our…
Consider the continuous-time Markov Branching Process. In critical case we consider a situation when the generating function of intensity of transformation of particles has the infinite second moment, but its tail regularly varies in sense…
In this article, we consider a stochastic linear quadratic control problem with partial observation. A near optimal control in the weak formulation is characterized. The main features of this paper are the presence of the control in the…
Reinforcement Learning Algorithms are predominantly developed for stationary environments, and the limited literature that considers nonstationary environments often involves specific assumptions about changes that can occur in transition…
We investigate constrained optimal control problems for linear stochastic dynamical systems evolving in discrete time. We consider minimization of an expected value cost over a finite horizon. Hard constraints are introduced first, and then…
We prove the existence of optimal strategies for agents with cumulative prospect theory preferences who trade in a continuous-time illiquid market, transcending known results which pertained only to risk-averse utility maximizers. The…
We study the problem of offline learning in automated decision systems under the contextual bandits model. We are given logged historical data consisting of contexts, (randomized) actions, and (nonnegative) rewards. A common goal is to…
This papers deals with the constrained discounted control of piecewise deterministic Markov process (PDMPs) in general Borel spaces. The control variable acts on the jump rate and transition measure, and the goal is to minimize the total…
We consider the optimal control of solutions of first order Hamilton-Jacobi equations, where the Hamiltonian is convex with linear growth. This models the problem of steering the propagation of a front by constructing an obstacle. We prove…
In this paper, we describe a novel approach to imitation learning that infers latent policies directly from state observations. We introduce a method that characterizes the causal effects of latent actions on observations while…
A new class of stochastic processes called independent and periodically identically distributed (i.p.i.d.) processes is defined to capture periodically varying statistical behavior. A novel Bayesian theory is developed for detecting a…
The proximal policy optimization (PPO) algorithm stands as one of the most prosperous methods in the field of reinforcement learning (RL). Despite its success, the theoretical understanding of PPO remains deficient. Specifically, it is…
We study decision timing problems on finite horizon with Poissonian information arrivals. In our model, a decision maker wishes to optimally time her action in order to maximize her expected reward. The reward depends on an unobservable…
Markov models are widely used to describe processes of stochastic dynamics. Here, we show that Markov models are a natural consequence of the dynamical principle of Maximum Caliber. First, we show that when there are different possible…
This paper is concerned with the deterministic optimal control of Ito stochastic systems with random coefficients. The necessary and sufficient conditions for the unique solvability of the optimal control problem with random coefficients…