English
Related papers

Related papers: Entropy-Regularized Certainty-Equivalent Bellman P…

200 papers

This paper develops a quantitative framework for analyzing the mean-square exponential stabilization of stochastic linear systems with multiplicative noise, focusing specifically on the optimal stabilizing rate, which characterizes the…

Optimization and Control · Mathematics 2025-12-15 Hui Jia , Yuan-Hua Ni , Guangchen Wang

Entropy regularized algorithms such as Soft Q-learning and Soft Actor-Critic, recently showed state-of-the-art performance on a number of challenging reinforcement learning (RL) tasks. The regularized formulation modifies the standard RL…

Machine Learning · Statistics 2019-10-15 Elena Smirnova , Elvis Dohmatob

In this paper, we consider scaling limits of exponential utility indifference prices for European contingent claims in the Bachelier model. We show that the scaling limit can be represented in terms of the \emph{specific relative entropy},…

Probability · Mathematics 2025-09-08 Yan Dolinksy , Xin Zhang

We study the policy iteration algorithm (PIA) for entropy-regularized stochastic control problems on an infinite time horizon with a large discount rate, focusing on two main scenarios. First, we analyze PIA with bounded coefficients where…

Optimization and Control · Mathematics 2025-05-28 Hung Vinh Tran , Zhenhua Wang , Yuming Paul Zhang

We address the problem of combined stochastic and impulse control for a market maker operating in a limit order book. The problem is formulated as a Hamilton-Jacobi-Bellman quasi-variational inequality (HJBQVI). We propose an implicit…

Mathematical Finance · Quantitative Finance 2025-12-25 Alexey Meteykin

Recently, \citet{SuttonMW15} introduced the emphatic temporal differences (ETD) algorithm for off-policy evaluation in Markov decision processes. In this short note, we show that the projected fixed-point equation that underlies ETD…

Machine Learning · Statistics 2015-08-25 Assaf Hallak , Aviv Tamar , Shie Mannor

There are no computationally feasible algorithms that provide solutions to the finite horizon Risk-sensitive Constrained Markov Decision Process (Risk-CMDP) problem, even for problems with moderate horizon. With an aim to design the same,…

Optimization and Control · Mathematics 2023-03-27 Vartika Singh , Veeraruna Kavitha

In this paper, we investigate optimal stopping problems in a continuous-time framework where only a discrete set of stopping dates is admissible, corresponding to the Bermudan option, within the so-called exploratory formulation. We…

Probability · Mathematics 2025-09-24 Noufel Frikha , Libo Li , Daniel Chee

While reinforcement learning has shown experimental success in a number of applications, it is known to be sensitive to noise and perturbations in the parameters of the system, leading to high variance in the total reward amongst different…

Systems and Control · Electrical Eng. & Systems 2024-12-02 Erfaun Noorani , Christos Mavridis , John Baras

A method for calculating multi-portfolio time consistent multivariate risk measures in discrete time is presented. Market models for $d$ assets with transaction costs or illiquidity and possible trading constraints are considered on a…

Risk Management · Quantitative Finance 2017-01-27 Zachary Feinstein , Birgit Rudloff

This work uses the entropy-regularised relaxed stochastic control perspective as a principled framework for designing reinforcement learning (RL) algorithms. Herein agent interacts with the environment by generating noisy controls…

Machine Learning · Computer Science 2023-09-18 Lukasz Szpruch , Tanut Treetanthiploet , Yufei Zhang

This paper presents a safe robust policy iteration (SR-PI) algorithm to design controllers with satisficing (good enough) performance and safety guarantee. This is in contrast to standard PI-based control design methods with no safety…

Systems and Control · Electrical Eng. & Systems 2020-09-16 Yuzhen Han , Hamidreza Modares

We study the problem of computing the value function from a discretely-observed trajectory of a continuous-time diffusion process. We develop a new class of algorithms based on easily implementable numerical schemes that are compatible with…

Machine Learning · Computer Science 2024-07-09 Wenlong Mou , Yuhua Zhu

We present an analysis of two thermodynamic techniques for determining equilibria of self-gravitating systems. One is the Lynden-Bell entropy maximization analysis that introduced violent relaxation. Since we do not use the Stirling…

Cosmology and Nongalactic Astrophysics · Physics 2015-06-03 Eric I. Barnes , Liliya L. R. Williams

This paper studies the mean-field Markov decision process (MDP) with the centralized stopping under the non-exponential discount. The problem differs fundamentally from most existing studies on mean-field optimal control/stopping due to its…

Optimization and Control · Mathematics 2025-01-22 Xiang Yu , Fengyi Yuan

Non-local correlations that obey the no-signalling principle contain intrinsic randomness. In particular, for a specific Bell experiment, one can derive relations between the amount of randomness produced, as quantified by the min-entropy…

Quantum Physics · Physics 2019-05-07 Boris Bourdoncle , Pei-Sheng Lin , Denis Rosset , Antonio Acín , Yeong-Cherng Liang

We consider a discrete-time version of the popular optimal dividend pay-out problem in risk theory. The novel aspect of our approach is that we allow for a risk averse insurer, i.e., instead of maximising the expected discounted dividends…

Probability · Mathematics 2015-12-02 Nicole Bäuerle , Anna Jaśkiewicz

In this paper, we derive rates of convergence in the high-dimensional central limit theorem for Polyak--Ruppert averaged iterates generated by entropy-regularized asynchronous Q-learning with linear function approximation and a polynomial…

Machine Learning · Statistics 2026-05-19 Artemy Rubtsov , Rahul Singh , Eric Moulines , Alexey Naumov , Sergey Samsonov

In this paper, we study the exponential utility indifference pricing of pure endowment policies within a stochastic-factor model for an insurer who also invests in a financial market. Our framework incorporates a hazard rate modeled as an…

Portfolio Management · Quantitative Finance 2025-07-30 Alessandra Cretarola , Benedetta Salterini

We derive an equality for non-equilibrium statistical mechanics in finite-dimensional quantum systems. The equality concerns the worst-case work output of a time-dependent Hamiltonian protocol in the presence of a Markovian heat bath. It…