Related papers: Policy Iteration for Exploratory Hamilton--Jacobi-…
This paper considers the problem of designing time-dependent, real-time control policies for controllable nonlinear diffusion processes, with the goal of obtaining maximally-informative observations about parameters of interest. More…
What happens when a continuously evolving stochastic process is interrupted with large changes at random intervals $\tau$ distributed as a power-law $\sim \tau^{-(1+\alpha)};\alpha>0$? Modeling the stochastic process by diffusion and the…
In this paper, we consider discrete-time infinite horizon problems of optimal control to a terminal set of states. These are the problems that are often taken as the starting point for adaptive dynamic programming. Under very general…
In this paper, we present a scalable deep learning approach to solve opinion dynamics stochastic optimal control problems with mean field term coupling in the dynamics and cost function. Our approach relies on the probabilistic…
We consider the infinite-horizon discounted optimal control problem formalized by Markov Decision Processes. We focus on several approximate variations of the Policy Iteration algorithm: Approximate Policy Iteration, Conservative Policy…
We study the exploratory Hamilton--Jacobi--Bellman (HJB) equation arising from the entropy-regularized exploratory control problem, which was formulated by Wang, Zariphopoulou and Zhou (J. Mach. Learn. Res., 21, 2020) in the context of…
We consider a stochastic control problem with the assumption that the system is controlled until the state process breaks the fixed barrier. Assuming some general conditions, it is proved that the resulting Hamilton Jacobi Bellman equations…
This paper studies an infinite horizon optimal control problem for discrete-time linear systems and quadratic criteria, both with random parameters which are independent and identically distributed with respect to time. A classical approach…
We investigated the unbounded diffusion observed in a time-dependent oval-shaped billiard and its suppression owing to inelastic collisions with the boundary. The main focus is on the behavior of the diffusion coefficient, which plays a key…
This paper revisits and extends the convergence and robustness properties of value and policy iteration algorithms for discrete-time linear quadratic regulator problems. In the model-based case, we extend current results concerning the…
We study a stochastic, continuous time model on a finite horizon for a firm that produces a single good. We model the production capacity as an Ito diffusion controlled by a nondecreasing process representing the cumulative investment. The…
We consider infinite horizon dynamic programming problems, where the control at each stage consists of several distinct decisions, each one made by one of several agents. In an earlier work we introduced a policy iteration algorithm, where…
In standard treatments of stochastic filtering one first has to estimate the values of the parameters of the model. Simply running the filter without considering the reliability of this estimate does not take into account this additional…
We introduce a regulated stochastic diffusion model for the recycling rate and formulate a joint control problem over production and process innovation via the dynamics of recycling investment and product pricing. The resulting stochastic…
This paper extends path integral control (PIC) to partially observed systems by formulating the problem in Gaussian belief space. PIC relies on the diffusion being proportional to the control channel -- the so-called matching condition --…
In this manuscript we consider a class optimal control problem for stochastic differential delay equations. First, we rewrite the problem in a suitable infinite-dimensional Hilbert space. Then, using the dynamic programming approach, we…
We consider the numerical solution of Hamilton-Jacobi-Bellman equations arising in stochastic control theory. We introduce a class of monotone approximation schemes relying on monotone interpolation. These schemes converge under very weak…
An advantageous feature of piecewise constant policy timestepping for Hamilton-Jacobi-Bellman (HJB) equations is that different linear approximation schemes, and indeed different meshes, can be used for the resulting linear equations for…
In this paper, we first conduct a study of the portfolio selection problem, incorporating both exogenous (proportional) and endogenous (resulting from liquidity risk, characterized by a stochastic process) transaction costs through the…
This paper studies an infinite horizon optimal control problem for discrete-time linear system and quadratic criteria, both with random parameters which are independent and identically distributed with respect to time. In this general…