Related papers: Entropy-Regularized Certainty-Equivalent Bellman P…
This work proposes new estimators for discrete optimal transport plans that enjoy Gaussian limits centered at the true solution. This behavior stands in stark contrast with the performance of existing estimators, including those based on…
An off policy reinforcement learning based control strategy is developed for the optimal tracking control problem to achieve the prescribed performance of full states during the learning process. The optimal tracking control problem is…
We present a novel continuous-time control strategy to exponentially stabilize an eigenstate of a Quantum Non-Demolition (QND) measurement operator. In open-loop, the system converges to a random eigenstate of the measurement operator. The…
In recent work it is shown that Q-learning with linear function approximation is stable, in the sense of bounded parameter estimates, under the $(\varepsilon,\kappa)$-tamed Gibbs policy; $\kappa$ is inverse temperature, and $\varepsilon>0$…
We study methods based on reproducing kernel Hilbert spaces for estimating the value function of an infinite-horizon discounted Markov reward process (MRP). We study a regularized form of the kernel least-squares temporal difference (LSTD)…
Entropy regularization has been widely used in policy optimization algorithms to enhance exploration and the robustness of the optimal control; however it also introduces an additional regularization bias. This work quantifies the impact of…
This paper studies the continuous-time reinforcement learning (RL) for optimal switching problems across multiple regimes. We consider a type of exploratory formulation under entropy regularization where the agent randomizes both the timing…
In this paper we study the robust invariant sets generation problem for discrete-time switched polynomial systems subject to disturbance inputs within the optimal control framework. A robust invariant set of interest is a set of states such…
This paper studies an optimal dividend problem for a company that aims to maximize the mean-variance (MV) objective of the accumulated discounted dividend payments up to its ruin time. The MV objective involves an integral form over a…
Forecasting accuracy is routinely optimised in financial prediction tasks even though investment and risk-management decisions are executed under transaction costs, market impact, capacity limits, and binding risk constraints. This paper…
In the paper, we consider the problem of robust approximation of transfer Koopman and Perron-Frobenius (P-F) operators from noisy time series data. In most applications, the time-series data obtained from simulation or experiment is…
This paper introduces a novel stochastic control framework to enhance the capabilities of automated investment managers, or robo-advisors, by accurately inferring clients' investment preferences from past activities. Our approach leverages…
We study a benchmarked risk-sensitive portfolio problem in a factor-based setting to bring together three strands of the literature: benchmarked risk-sensitive investment management, the Kuroda-Nagai change-of-measure method, and the free…
Learning and optimal control under robust Markov decision processes (MDPs) have received increasing attention, yet most existing theory, algorithms, and applications focus on finite-horizon or discounted models. Long-run average-reward…
This study investigates an optimal investment problem for an insurance company operating under the Cramer-Lundberg risk model, where investments are made in both a risky asset and a risk-free asset. In contrast to other literature that…
We consider the problem of learning the Hamiltonian of a quantum system from estimates of Gibbs-state expectation values. Various methods for achieving this task were proposed recently, both from a practical and theoretical point of view.…
For job scheduling systems, where jobs require some amount of processing and then leave the system, it is natural for each user to provide an estimate of their job's time requirement in order to aid the scheduler. However, if there is no…
This paper studies the dividend and capital injection problem under a diffusion risk model with general discount functions. A proportional cost is imposed when injecting capitals. For exponential discounting as time-consistent benchmark, we…
We study the exploratory Hamilton--Jacobi--Bellman (HJB) equation arising from the entropy-regularized exploratory control problem, which was formulated by Wang, Zariphopoulou and Zhou (J. Mach. Learn. Res., 21, 2020) in the context of…
Sample-based trajectory optimisers are a promising tool for the control of robotics with non-differentiable dynamics and cost functions. Contemporary approaches derive from a restricted subclass of stochastic optimal control where the…