English
Related papers

Related papers: A variational formula for risk-sensitive reward

200 papers

We consider continuous-time stochastic optimal control problems featuring Conditional Value-at-Risk (CVaR) in the objective. The major difficulty in these problems arises from time-inconsistency, which prevents us from directly using…

Optimization and Control · Mathematics 2020-05-27 Christopher W. Miller , Insoon Yang

In the paper average reward per unit time and average risk sensitive reward functionals are considered for controlled nonhomogeneous Markov processes. Existence of solutions to suitable Bellman equations is shown. Continuity of the value…

Optimization and Control · Mathematics 2025-06-19 Łukasz Stettner

The problem of reward design examines the interaction between a leader and a follower, where the leader aims to shape the follower's behavior to maximize the leader's payoff by modifying the follower's reward function. Current approaches to…

Optimization and Control · Mathematics 2024-06-10 Shuo Wu , Haoxiang Ma , Jie Fu , Shuo Han

Existing work on risk-sensitive reinforcement learning - both for symmetric and downside risk measures - has typically used direct Monte-Carlo estimation of policy gradients. While this approach yields unbiased gradient estimates, it also…

Machine Learning · Computer Science 2020-07-09 Thomas Spooner , Rahul Savani

This article is related to risk-sensitive nonzero-sum stochastic differential games in the Markovian framework. This game takes into account the attitudes of the players toward risk and the utility is of exponential form. We show the…

Optimization and Control · Mathematics 2014-12-04 Said Hamadène , Rui Mu

We consider an optimal control problem with ergodic (long term average) reward for a McKean-Vlasov dynamics, where the coefficients of a controlled stochastic differential equation depend on the marginal law of the solution. Starting from…

Optimization and Control · Mathematics 2025-11-25 Marco Fuhrman , Silvia Rudà

Sharpe ratio (also known as reward-to-variability ratio) is a widely-used metric in finance, which measures the additional return at the cost of per unit of increased risk (standard deviation of return). However, the optimization of Sharpe…

Artificial Intelligence · Computer Science 2025-09-03 Shuai Ma , Guangwu Liu , Li Xia

This paper addresses the inverse optimal control problem of finding the state weighting function that leads to a quadratic value function when the cost on the input is fixed to be quadratic. The paper focuses on a class of infinite horizon…

Optimization and Control · Mathematics 2022-11-21 Luis Rodrigues

In this paper, we establish a general stochastic maximum principle for optimal control for systems described by a continuous-time Markov regime-switching stochastic recursive utilities model. The control domain is postulated not to be…

Optimization and Control · Mathematics 2019-05-02 Liangquan Zhang , Xun Li

This paper is concerned with an optimal reinsurance and investment problem for an insurance firm under the criterion of mean-variance. The driving Brownian motion and the rate in return of the risky asset price dynamic equation cannot be…

Optimization and Control · Mathematics 2020-06-04 Shihao Zhu , Jingtao Shi

We obtain the variational equations for backward stochastic differential equations in recursive stochastic optimal control problems, and then get the maximum principle which is novel. The control domain need not be convex, and the generator…

Optimization and Control · Mathematics 2015-07-14 Mingshang Hu

We consider a dynamic portfolio optimization problem that incorporates predictable returns, instantaneous transaction costs, price impact, and stochastic volatility, extending the classical results of Garleanu and Pedersen (2013), which…

Computational Finance · Quantitative Finance 2025-07-24 Patrick Chan , Ronnie Sircar , Iosif Zimbidis

This paper deals with discrete-time Markov control processes on a general state space. A long-run risk-sensitive average cost criterion is used as a performance measure. The one-step cost function is nonnegative and possibly unbounded.…

Risk Management · Quantitative Finance 2016-08-14 Anna Jaśkiewicz

Inference-time scaling methods rely on Process Reward Models (PRMs), which are often poorly calibrated and overestimate success probabilities. We propose, to our knowledge, the first use of conditional optimal transport for calibrating…

Machine Learning · Computer Science 2026-05-13 Rachel Ma , Dylan Hadfield-Menell , Kristjan Greenewald

This paper is devoted to solving a time-inconsistent risk-sensitive control problem with parameter $\e$ and its limit case ($\e\rightarrow0^+$) for countable-stated Markov decision processes (MDPs for short). Since the cost functional is…

Optimization and Control · Mathematics 2020-10-22 Hongwei Mei

This paper concerns an optimal dividend distribution problem for an insurance company with surplus-dependent premium. In the absence of dividend payments, such a risk process is a particular case of so-called piecewise deterministic Markov…

Portfolio Management · Quantitative Finance 2016-04-26 Ewa Marciniak , Zbigniew Palmowski

Motivated by the design of fast reinforcement learning algorithms, we study the diffusive limit of a class of pure jump ergodic stochastic control problems. We show that, whenever the intensity of jumps is large enough, the approximation…

Optimization and Control · Mathematics 2022-10-03 Marc Abeille , Bruno Bouchard , Lorenzo Croissant

We consider the optimal stopping problem consisting in, given a strong Markov process, a reward function and a discount rate, finding the stopping time such that the expected reward at the stopping time is maximum. The approach we follow,…

Probability · Mathematics 2014-05-30 Fabián Crocce

This paper studies an optimal investment-reinsurance problem for an insurer (she) under the Cram\'er--Lundberg model with monotone mean--variance (MMV) criterion. At any time, the insurer can purchase reinsurance (or acquire new business)…

Portfolio Management · Quantitative Finance 2024-05-30 Xiaomin Shi , Zuo Quan Xu

We present a branch-and-bound algorithm for globally solving parabolic optimal control problems with binary switches that have bounded variation and possibly need to satisfy further combinatorial constraints. More precisely, for a given…

Optimization and Control · Mathematics 2024-01-19 Christoph Buchheim , Alexandra Grütering , Christian Meyer