中文
相关论文

相关论文: A variational formula for risk-sensitive reward

200 篇论文

We study the problem of reinforcement learning in infinite-horizon discounted linear Markov decision processes (MDPs), and propose the first computationally efficient algorithm achieving rate-optimal regret guarantees in this setting. Our…

机器学习 · 计算机科学 2026-03-16 Antoine Moulin , Gergely Neu , Luca Viano

This paper studies the optimization of Markov decision processes (MDPs) from a risk-seeking perspective, where the risk is measured by conditional value-at-risk (CVaR). The objective is to find a policy that maximizes the long-run CVaR of…

最优化与控制 · 数学 2023-12-05 Li Xia , Zhihui Yu , Peter W. Glynn

We analyse an optimal control with the following features: the dynamical system is linear, and the dependence upon the control parameter is affine. More precisely we consider $\dot x_\alpha(t) = (G + \alpha(t) F)x_\alpha(t)$, where $G$ and…

偏微分方程分析 · 数学 2012-03-26 Vincent Calvez , Pierre Gabriel

We consider reinforcement learning in changing Markov Decision Processes where both the state-transition probabilities and the reward functions may vary over time. For this problem setting, we propose an algorithm using a sliding window…

机器学习 · 计算机科学 2018-05-28 Pratik Gajane , Ronald Ortner , Peter Auer

In this paper we extend temporal difference policy evaluation algorithms to performance criteria that include the variance of the cumulative reward. Such criteria are useful for risk management, and are important in domains such as finance…

机器学习 · 计算机科学 2013-10-15 Aviv Tamar , Dotan Di Castro , Shie Mannor

We study optimal investment strategies that maximize expected utility from consumption and terminal wealth in a pure-jump asset price model with Markov-modulated (regime switching) jump-size distributions. We give sufficient conditions for…

投资组合管理 · 定量金融 2014-06-13 Oscar Lopez , Rafael Serrano

We study the problem of optimal dividend payout from a surplus process governed by Brownian motion with drift under the additional constraint of ratcheting, i.e. the dividend rate can never decrease. We solve the resulting two-dimensional…

概率论 · 数学 2020-12-22 Hansjoerg Albrecher , Pablo Azcue , Nora Muler

In this paper we consider a variation of the Merton's problem with added stochastic volatility and finite time horizon. It is known that the corresponding optimal control problem may be reduced to a linear parabolic boundary problem under…

数理金融 · 定量金融 2015-05-28 Elena Boguslavskaya , Dmitry Muravey

We consider a liquidation problem in which a risk-averse trader tries to liquidate a fixed quantity of an asset in the presence of market impact and random price fluctuations. The trader encounters a trade-off between the transaction costs…

交易与市场微观结构 · 定量金融 2022-01-31 Seungki Min , Ciamac C. Moallemi , Costis Maglaras

Efficient exploration remains a challenging problem in reinforcement learning, especially for tasks where extrinsic rewards from environments are sparse or even totally disregarded. Significant advances based on intrinsic motivation show…

机器学习 · 计算机科学 2024-04-03 Chenjia Bai , Peng Liu , Kaiyu Liu , Lingxiao Wang , Yingnan Zhao , Lei Han

We consider an optimal investment and risk control problem for an insurer under the mean-variance (MV) criterion. By introducing a deterministic auxiliary process defined forward in time, we formulate an alternative time-consistent problem…

投资组合管理 · 定量金融 2021-01-12 Yang Shen , Bin Zou

We propose a numerical recipe for risk evaluation defined by a backward stochastic differential equation. Using dual representation of the risk measure, we convert the risk valuation to a stochastic control problem where the control is a…

最优化与控制 · 数学 2020-08-24 Andrzej Ruszczynski , Jianing Yao

We study optimal investment problem for a diffusion market consisting of a finite number of risky assets (for example, bonds, stocks and options). Risky assets evolution is described by Ito's equation, and the number of risky assets can be…

概率论 · 数学 2008-12-02 Nikolai Dokuchaev

We consider the optimal control problem for a linear conditional McKean-Vlasov equation with quadratic cost functional. The coefficients of the system and the weigh-ting matrices in the cost functional are allowed to be adapted processes…

概率论 · 数学 2017-03-09 Huyên Pham

We study a first-order primal-dual subgradient method to optimize risk-constrained risk-penalized optimization problems, where risk is modeled via the popular conditional value at risk (CVaR) measure. The algorithm processes independent and…

最优化与控制 · 数学 2021-09-03 Avinash N. Madavan , Subhonmesh Bose

This article studies a portfolio optimization problem, where the market consisting of several stocks is modeled by a multi-dimensional jump-diffusion process with age-dependent semi-Markov modulated coefficients. We study risk sensitive…

投资组合管理 · 定量金融 2019-10-21 Milan Kumar Das , Anindya Goswami , Nimit Rana

Overconservatism has long been recognized as a major issue with robust optimization, despite its key advantages of tractability, performance guarantee, and limited information. To address this issue, a new criterion is proposed that can…

最优化与控制 · 数学 2026-03-20 Yingjie Lan

We consider an infinite horizon discounted optimal control problem for piecewise deterministic Markov processes, where a piecewise open-loop control acts continuously on the jump dynamics and on the deterministic flow. For this class of…

最优化与控制 · 数学 2015-12-08 Elena Bandini

Many popular practical reinforcement learning (RL) algorithms employ evolving reward functions-through techniques such as reward shaping, entropy regularization, or curriculum learning-yet their theoretical foundations remain…

机器学习 · 计算机科学 2025-10-15 Rui Hu , Yu Chen , Longbo Huang

We consider undiscounted reinforcement learning in Markov decision processes (MDPs) where both the reward functions and the state-transition probabilities may vary (gradually or abruptly) over time. For this problem setting, we propose an…

机器学习 · 计算机科学 2019-09-11 Pratik Gajane , Ronald Ortner , Peter Auer