中文
相关论文

相关论文: Boosting CVaR Policy Optimization with Quantile Gr…

200 篇论文

In this paper a class of combinatorial optimization problems is discussed. It is assumed that a solution can be constructed in two stages. The current first-stage costs are precisely known, while the future second-stage costs are only known…

数据结构与算法 · 计算机科学 2018-12-20 Marc Goerigk , Adam Kasperski , Pawel Zielinski

Optimizing static risk-averse objectives in Markov decision processes is difficult because they do not admit standard dynamic programming equations common in Reinforcement Learning (RL) algorithms. Dynamic programming decompositions that…

最优化与控制 · 数学 2024-07-04 Jia Lin Hau , Erick Delage , Mohammad Ghavamzadeh , Marek Petrik

This paper considers variational inequalities (VI) defined by the conditional value-at-risk (CVaR) of uncertain functions and provides three stochastic approximation schemes to solve them. All methods use an empirical estimate of the CVaR…

最优化与控制 · 数学 2022-11-16 Jasper Verbree , Ashish Cherukuri

Motivated by the poor performance of cross-validation in settings where data are scarce, we propose a novel estimator of the out-of-sample performance of a policy in data-driven optimization.Our approach exploits the optimization problem's…

最优化与控制 · 数学 2022-08-04 Vishal Gupta , Michael Huang , Paat Rusmevichientong

This paper develops a safety analysis method for stochastic systems that is sensitive to the possibility and severity of rare harmful outcomes. We define risk-sensitive safe sets as sub-level sets of the solution to a non-standard optimal…

系统与控制 · 电气工程与系统科学 2022-06-28 Margaret P. Chapman , Riccardo Bonalli , Kevin M. Smith , Insoon Yang , Marco Pavone , Claire J. Tomlin

This paper proposes a safety analysis method that facilitates a tunable balance between the worst-case and risk-neutral perspectives. First, we define a risk-sensitive safe set to specify the degree of safety attained by a stochastic…

系统与控制 · 电气工程与系统科学 2020-07-28 Margaret P. Chapman , Jonathan P. Lacotte , Kevin M. Smith , Insoon Yang , Yuxi Han , Marco Pavone , Claire J. Tomlin

In many sequential decision-making problems one is interested in minimizing an expected cumulative cost while taking into account \emph{risk}, i.e., increased awareness of events of small probability and high consequences. Accordingly, the…

人工智能 · 计算机科学 2017-04-07 Yinlam Chow , Mohammad Ghavamzadeh , Lucas Janson , Marco Pavone

A method for quantile-based, semi-parametric historical simulation estimation of multiple step ahead Value-at-Risk (VaR) and Expected Shortfall (ES) models is developed. It uses the quantile loss function, analogous to how the…

统计金融 · 定量金融 2025-03-06 Richard Gerlach , Antonio Naimoli , Giuseppe Storti

Quantile regression (QR) relies on the estimation of conditional quantiles and explores the relationships between independent and dependent variables. At high probability levels, classical QR methods face extrapolation difficulties due to…

We consider the portfolio optimization with risk measured by conditional value-at-risk, based on the stress event of chosen asset being equal to the opposite of its value-at-risk level, under the normality assumption. Solvability conditions…

最优化与控制 · 数学 2017-03-07 Anna Zalewska

Policy gradient based reinforcement learning algorithms coupled with neural networks have shown success in learning complex policies in the model free continuous action space control setting. However, explicitly parameterized policies are…

机器学习 · 计算机科学 2019-09-30 Oliver Richter , Roger Wattenhofer

We propose a method for finding approximate compilations of quantum unitary transformations, based on techniques from policy gradient reinforcement learning. The choice of a stochastic policy allows us to rephrase the optimization problem…

量子物理 · 物理学 2022-09-14 David A. Herrera-Martí

We consider a class of chance-constrained programs in which profit needs to be maximized while enforcing that a given adverse event remains rare. Using techniques from large deviations and extreme value theory, we show how the optimal value…

最优化与控制 · 数学 2025-11-12 Jose Blanchet , Joost Jorritsma , Bert Zwart

Conditional value-at-risk (CoVaR) is one of the most important measures of systemic risk. It is defined as the high quantile conditional on a related variable being extreme, widely used in the field of quantitative risk management. In this…

统计方法学 · 统计学 2026-02-12 Zhaowen Wang , Yutao Liu , Deyuan Li

The multi-armed bandit (MAB) problem is a ubiquitous decision-making problem that exemplifies the exploration-exploitation tradeoff. Standard formulations exclude risk in decision making. Risk notably complicates the basic reward-maximising…

机器学习 · 计算机科学 2021-02-05 Joel Q. L. Chang , Qiuyu Zhu , Vincent Y. F. Tan

This paper investigates the use of prior computation to estimate the value function to improve sample efficiency in on-policy policy gradient methods in reinforcement learning. Our approach is to estimate the value function from prior…

机器学习 · 计算机科学 2023-02-06 Md Masudur Rahman , Yexiang Xue

We consider optimal allocation problems with Conditional Value-At-Risk (CVaR) constraint. We prove, under very mild assumptions, the convergence of the Sample Average Approximation method (SAA) applied to this problem, and we also exhibit a…

投资组合管理 · 定量金融 2025-05-19 Jérôme Lelong , Véronique Maume-Deschamps , William Thevenot

Microgrid operation is highly vulnerable to short-term load uncertainty, while conventional predict-then-optimize pipelines cannot fully align probabilistic forecasting quality with downstream robust scheduling performance. This paper…

系统与控制 · 电气工程与系统科学 2026-04-21 Tingwei Cao , Yan Xu

In this study, we address the challenge of portfolio optimization, a critical aspect of managing investment risks and maximizing returns. The mean-CVaR portfolio is considered a promising method due to today's unstable financial market…

投资组合管理 · 定量金融 2023-09-22 Kei Nakagawa , Masaya Abe , Seiichi Kuroki

The control variates (CV) method is widely used in policy gradient estimation to reduce the variance of the gradient estimators in practice. A control variate is applied by subtracting a baseline function from the state-action value…

机器学习 · 计算机科学 2021-08-12 Yuanyi Zhong , Yuan Zhou , Jian Peng