English
Related papers

Related papers: Entropy-Regularized Certainty-Equivalent Bellman P…

200 papers

This work proposes new estimators for discrete optimal transport plans that enjoy Gaussian limits centered at the true solution. This behavior stands in stark contrast with the performance of existing estimators, including those based on…

Statistics Theory · Mathematics 2025-05-08 Shuyu Liu , Florentina Bunea , Jonathan Niles-Weed

An off policy reinforcement learning based control strategy is developed for the optimal tracking control problem to achieve the prescribed performance of full states during the learning process. The optimal tracking control problem is…

Systems and Control · Electrical Eng. & Systems 2020-09-02 C. Li , Y. Wang , F. Liu , M. Buss

We present a novel continuous-time control strategy to exponentially stabilize an eigenstate of a Quantum Non-Demolition (QND) measurement operator. In open-loop, the system converges to a random eigenstate of the measurement operator. The…

Quantum Physics · Physics 2019-06-19 Gerardo Cardona , Alain Sarlette , Pierre Rouchon

In recent work it is shown that Q-learning with linear function approximation is stable, in the sense of bounded parameter estimates, under the $(\varepsilon,\kappa)$-tamed Gibbs policy; $\kappa$ is inverse temperature, and $\varepsilon>0$…

Machine Learning · Computer Science 2026-02-09 Prashant Mehta , Sean Meyn

We study methods based on reproducing kernel Hilbert spaces for estimating the value function of an infinite-horizon discounted Markov reward process (MRP). We study a regularized form of the kernel least-squares temporal difference (LSTD)…

Machine Learning · Statistics 2021-09-27 Yaqi Duan , Mengdi Wang , Martin J. Wainwright

Entropy regularization has been widely used in policy optimization algorithms to enhance exploration and the robustness of the optimal control; however it also introduces an additional regularization bias. This work quantifies the impact of…

Optimization and Control · Mathematics 2025-03-25 Deven Sethi , David Šiška , Yufei Zhang

This paper studies the continuous-time reinforcement learning (RL) for optimal switching problems across multiple regimes. We consider a type of exploratory formulation under entropy regularization where the agent randomizes both the timing…

Optimization and Control · Mathematics 2025-12-23 Yijie Huang , Mengge Li , Xiang Yu , Zhou Zhou

In this paper we study the robust invariant sets generation problem for discrete-time switched polynomial systems subject to disturbance inputs within the optimal control framework. A robust invariant set of interest is a set of states such…

Discrete Mathematics · Computer Science 2021-03-22 Bai Xue , Naijun Zhan

This paper studies an optimal dividend problem for a company that aims to maximize the mean-variance (MV) objective of the accumulated discounted dividend payments up to its ruin time. The MV objective involves an integral form over a…

Optimization and Control · Mathematics 2025-08-19 Jingyi Cao , Dongchen Li , Virginia R. Young , Bin Zou

Forecasting accuracy is routinely optimised in financial prediction tasks even though investment and risk-management decisions are executed under transaction costs, market impact, capacity limits, and binding risk constraints. This paper…

Econometrics · Economics 2026-01-14 Craig S Wright

In the paper, we consider the problem of robust approximation of transfer Koopman and Perron-Frobenius (P-F) operators from noisy time series data. In most applications, the time-series data obtained from simulation or experiment is…

Optimization and Control · Mathematics 2020-01-08 Subhrajit Sinha , Huang Bowen , Umesh Vaidya

This paper introduces a novel stochastic control framework to enhance the capabilities of automated investment managers, or robo-advisors, by accurately inferring clients' investment preferences from past activities. Our approach leverages…

Optimization and Control · Mathematics 2024-06-05 Haoyang Cao , Zhengqi Wu , Renyuan Xu

We study a benchmarked risk-sensitive portfolio problem in a factor-based setting to bring together three strands of the literature: benchmarked risk-sensitive investment management, the Kuroda-Nagai change-of-measure method, and the free…

Portfolio Management · Quantitative Finance 2026-04-28 Sebastien Lleo , Wolfgang Runggaldier

Learning and optimal control under robust Markov decision processes (MDPs) have received increasing attention, yet most existing theory, algorithms, and applications focus on finite-horizon or discounted models. Long-run average-reward…

Optimization and Control · Mathematics 2025-12-12 Shengbo Wang , Nian Si

This study investigates an optimal investment problem for an insurance company operating under the Cramer-Lundberg risk model, where investments are made in both a risky asset and a risk-free asset. In contrast to other literature that…

Mathematical Finance · Quantitative Finance 2024-06-25 J. Cerda-Hernandez , A. Sikov , A. Ramos

We consider the problem of learning the Hamiltonian of a quantum system from estimates of Gibbs-state expectation values. Various methods for achieving this task were proposed recently, both from a practical and theoretical point of view.…

Quantum Physics · Physics 2024-10-31 Adam Artymowicz , Hamza Fawzi , Omar Fawzi , Samuel O. Scalet

For job scheduling systems, where jobs require some amount of processing and then leave the system, it is natural for each user to provide an estimate of their job's time requirement in order to aid the scheduler. However, if there is no…

Computer Science and Game Theory · Computer Science 2022-02-14 Isaac Grosof , Michael Mitzenmacher

This paper studies the dividend and capital injection problem under a diffusion risk model with general discount functions. A proportional cost is imposed when injecting capitals. For exponential discounting as time-consistent benchmark, we…

Mathematical Finance · Quantitative Finance 2025-05-30 Sang Hu , Zihan Zhou

We study the exploratory Hamilton--Jacobi--Bellman (HJB) equation arising from the entropy-regularized exploratory control problem, which was formulated by Wang, Zariphopoulou and Zhou (J. Mach. Learn. Res., 21, 2020) in the context of…

Optimization and Control · Mathematics 2021-09-22 Wenpin Tang , Paul Yuming Zhang , Xun Yu Zhou

Sample-based trajectory optimisers are a promising tool for the control of robotics with non-differentiable dynamics and cost functions. Contemporary approaches derive from a restricted subclass of stochastic optimal control where the…

Robotics · Computer Science 2021-10-07 Tom Lefebvre , Guillaume Crevecoeur