English
Related papers

Related papers: A variational formula for risk-sensitive reward

200 papers

We study the problem of reinforcement learning in infinite-horizon discounted linear Markov decision processes (MDPs), and propose the first computationally efficient algorithm achieving rate-optimal regret guarantees in this setting. Our…

Machine Learning · Computer Science 2026-03-16 Antoine Moulin , Gergely Neu , Luca Viano

This paper studies the optimization of Markov decision processes (MDPs) from a risk-seeking perspective, where the risk is measured by conditional value-at-risk (CVaR). The objective is to find a policy that maximizes the long-run CVaR of…

Optimization and Control · Mathematics 2023-12-05 Li Xia , Zhihui Yu , Peter W. Glynn

We analyse an optimal control with the following features: the dynamical system is linear, and the dependence upon the control parameter is affine. More precisely we consider $\dot x_\alpha(t) = (G + \alpha(t) F)x_\alpha(t)$, where $G$ and…

Analysis of PDEs · Mathematics 2012-03-26 Vincent Calvez , Pierre Gabriel

We consider reinforcement learning in changing Markov Decision Processes where both the state-transition probabilities and the reward functions may vary over time. For this problem setting, we propose an algorithm using a sliding window…

Machine Learning · Computer Science 2018-05-28 Pratik Gajane , Ronald Ortner , Peter Auer

In this paper we extend temporal difference policy evaluation algorithms to performance criteria that include the variance of the cumulative reward. Such criteria are useful for risk management, and are important in domains such as finance…

Machine Learning · Computer Science 2013-10-15 Aviv Tamar , Dotan Di Castro , Shie Mannor

We study optimal investment strategies that maximize expected utility from consumption and terminal wealth in a pure-jump asset price model with Markov-modulated (regime switching) jump-size distributions. We give sufficient conditions for…

Portfolio Management · Quantitative Finance 2014-06-13 Oscar Lopez , Rafael Serrano

We study the problem of optimal dividend payout from a surplus process governed by Brownian motion with drift under the additional constraint of ratcheting, i.e. the dividend rate can never decrease. We solve the resulting two-dimensional…

Probability · Mathematics 2020-12-22 Hansjoerg Albrecher , Pablo Azcue , Nora Muler

In this paper we consider a variation of the Merton's problem with added stochastic volatility and finite time horizon. It is known that the corresponding optimal control problem may be reduced to a linear parabolic boundary problem under…

Mathematical Finance · Quantitative Finance 2015-05-28 Elena Boguslavskaya , Dmitry Muravey

We consider a liquidation problem in which a risk-averse trader tries to liquidate a fixed quantity of an asset in the presence of market impact and random price fluctuations. The trader encounters a trade-off between the transaction costs…

Trading and Market Microstructure · Quantitative Finance 2022-01-31 Seungki Min , Ciamac C. Moallemi , Costis Maglaras

Efficient exploration remains a challenging problem in reinforcement learning, especially for tasks where extrinsic rewards from environments are sparse or even totally disregarded. Significant advances based on intrinsic motivation show…

Machine Learning · Computer Science 2024-04-03 Chenjia Bai , Peng Liu , Kaiyu Liu , Lingxiao Wang , Yingnan Zhao , Lei Han

We consider an optimal investment and risk control problem for an insurer under the mean-variance (MV) criterion. By introducing a deterministic auxiliary process defined forward in time, we formulate an alternative time-consistent problem…

Portfolio Management · Quantitative Finance 2021-01-12 Yang Shen , Bin Zou

We propose a numerical recipe for risk evaluation defined by a backward stochastic differential equation. Using dual representation of the risk measure, we convert the risk valuation to a stochastic control problem where the control is a…

Optimization and Control · Mathematics 2020-08-24 Andrzej Ruszczynski , Jianing Yao

We study optimal investment problem for a diffusion market consisting of a finite number of risky assets (for example, bonds, stocks and options). Risky assets evolution is described by Ito's equation, and the number of risky assets can be…

Probability · Mathematics 2008-12-02 Nikolai Dokuchaev

We consider the optimal control problem for a linear conditional McKean-Vlasov equation with quadratic cost functional. The coefficients of the system and the weigh-ting matrices in the cost functional are allowed to be adapted processes…

Probability · Mathematics 2017-03-09 Huyên Pham

We study a first-order primal-dual subgradient method to optimize risk-constrained risk-penalized optimization problems, where risk is modeled via the popular conditional value at risk (CVaR) measure. The algorithm processes independent and…

Optimization and Control · Mathematics 2021-09-03 Avinash N. Madavan , Subhonmesh Bose

This article studies a portfolio optimization problem, where the market consisting of several stocks is modeled by a multi-dimensional jump-diffusion process with age-dependent semi-Markov modulated coefficients. We study risk sensitive…

Portfolio Management · Quantitative Finance 2019-10-21 Milan Kumar Das , Anindya Goswami , Nimit Rana

Overconservatism has long been recognized as a major issue with robust optimization, despite its key advantages of tractability, performance guarantee, and limited information. To address this issue, a new criterion is proposed that can…

Optimization and Control · Mathematics 2026-03-20 Yingjie Lan

We consider an infinite horizon discounted optimal control problem for piecewise deterministic Markov processes, where a piecewise open-loop control acts continuously on the jump dynamics and on the deterministic flow. For this class of…

Optimization and Control · Mathematics 2015-12-08 Elena Bandini

Many popular practical reinforcement learning (RL) algorithms employ evolving reward functions-through techniques such as reward shaping, entropy regularization, or curriculum learning-yet their theoretical foundations remain…

Machine Learning · Computer Science 2025-10-15 Rui Hu , Yu Chen , Longbo Huang

We consider undiscounted reinforcement learning in Markov decision processes (MDPs) where both the reward functions and the state-transition probabilities may vary (gradually or abruptly) over time. For this problem setting, we propose an…

Machine Learning · Computer Science 2019-09-11 Pratik Gajane , Ronald Ortner , Peter Auer