English
Related papers

Related papers: Risk-Averse Trust Region Optimization for Reward-V…

200 papers

Risk-averse Constrained Reinforcement Learning (RaCRL) aims to learn policies that minimise the likelihood of rare and catastrophic constraint violations caused by an environment's inherent randomness. In general, risk-aversion leads to…

Machine Learning · Computer Science 2025-08-28 James McCarthy , Radu Marinescu , Elizabeth Daly , Ivana Dusparic

A risk measure that is consistent with the second-order stochastic dominance and additive for sums of independent random variables can be represented as a weighted entropic risk measure (WERM). The expected utility maximization problem with…

Mathematical Finance · Quantitative Finance 2021-12-07 Jianming Xia

Traditional imitation learning provides a set of methods and algorithms to learn a reward function or policy from expert demonstrations. Learning from demonstration has been shown to be advantageous for navigation tasks as it allows for…

Robotics · Computer Science 2021-08-03 Christian Ellis , Maggie Wigness , John G. Rogers , Craig Lennon , Lance Fiondella

This paper develops an online inverse reinforcement learning algorithm aimed at efficiently recovering a reward function from ongoing observations of an agent's actions. To reduce the computation time and storage space in reward estimation,…

Robotics · Computer Science 2017-08-01 Kun Li , Joel W. Burdick

This paper discusses an alternative explanation for the empirical findings contradicting the positive relationship between risk (variance) and reward (expected return). We show that these contradicting results might be due to the false…

Risk Management · Quantitative Finance 2017-04-19 Mihaly Ormos , Dusan Timotity

Policy gradient methods, which have been extensively studied in the last decade, offer an effective and efficient framework for reinforcement learning problems. However, their performances can often be unsatisfactory, suffering from…

Machine Learning · Computer Science 2026-01-27 Shihab Ahmed , El Houcine Bergou , Aritra Dutta , Yue Wang

Safe reinforcement learning (RL) seeks to mitigate unsafe behaviors that arise from exploration during training by reducing constraint violations while maintaining task performance. Existing approaches typically rely on a single policy to…

Robotics · Computer Science 2026-05-12 Murad Dawood , Usama Ahmed Siddiquie , Shahram Khorshidi , Maren Bennewitz

Rewards serve as a measure of user satisfaction and act as a limiting factor in interactive recommender systems. In this research, we focus on the problem of learning to reward (LTR), which is fundamental to reinforcement learning. Previous…

Machine Learning · Computer Science 2023-10-31 Jialin Liu , Xinyan Su , Zeyu He , Xiangyu Zhao , Jun Li

This paper proposes a suite of rationality measures and associated theory for reinforcement learning agents, a property increasingly critical yet rarely explored. We define an action in deployment to be perfectly rational if it maximises…

Machine Learning · Computer Science 2026-05-05 Kejiang Qian , Amos Storkey , Fengxiang He

The investor is interested in the expected return and he is also concerned about the risk and the uncertainty assumed by the investment. One of the most popular concepts used to measure the risk and the uncertainty is the variance and/or…

Statistical Finance · Quantitative Finance 2008-12-02 Andreia Dionisio , Rui Menezes , Diana A. Mendes

In a typical stochastic multi-armed bandit problem, the objective is often to maximize the expected sum of rewards over some time horizon $T$. While the choice of a strategy that accomplishes that is optimal with no additional information,…

Machine Learning · Computer Science 2023-11-01 Reda Alami , Mohammed Mahfoud , Mastane Achab

In risk-averse reinforcement learning (RL), the goal is to optimize some risk measure of the returns. A risk measure often focuses on the worst returns out of the agent's experience. As a result, standard methods for risk-averse RL often…

Machine Learning · Computer Science 2022-10-13 Ido Greenberg , Yinlam Chow , Mohammad Ghavamzadeh , Shie Mannor

Most work in mechanism design assumes that buyers are risk neutral; some considers risk aversion arising due to a non-linear utility for money. Yet behavioral studies have established that real agents exhibit risk attitudes which cannot be…

Computer Science and Game Theory · Computer Science 2018-03-13 Shuchi Chawla , Kira Goldner , J. Benjamin Miller , Emmanouil Pountourakis

In this paper, we study properties of certain risk measures associated with acceptance sets. These sets describe regulatory preconditions that have to be fulfilled by financial institutions to pass a given acceptance test. If the financial…

Optimization and Control · Mathematics 2021-10-07 Marcel Marohn , Christiane Tammer

Continued interest in sustainable investing calls for an axiomatic approach to measures of risk and reward that focus not only on financial returns, but also on measures of environmental and social sustainability, i.e. environmental,…

Mathematical Finance · Quantitative Finance 2026-02-19 Gabriele Torri , Rosella Giacometti , Darinka Dentcheva , Svetlozar T. Rachev , W. Brent Lindquist

Multi-turn tool calling is challenging for Large Language Models (LLMs) because rewards are sparse and exploration is expensive. A common recipe, SFT followed by GRPO, can stall when within-group reward variation is low (e.g., more rollouts…

Artificial Intelligence · Computer Science 2026-02-04 Haitian Zhong , Jixiu Zhai , Lei Song , Jiang Bian , Qiang Liu , Tieniu Tan

We consider a Markov decision process subject to model uncertainty in a Bayesian framework, where we assume that the state process is observed but its law is unknown to the observer. In addition, while the state process and the controls are…

Optimization and Control · Mathematics 2022-06-22 Tomasz R. Bielecki , Igor Cialenco , Andrzej Ruszczyński

In safety-critical applications of reinforcement learning such as healthcare and robotics, it is often desirable to optimize risk-sensitive objectives that account for tail outcomes rather than expected reward. We prove the first regret…

Machine Learning · Computer Science 2022-10-12 O. Bastani , Y. J. Ma , E. Shen , W. Xu

We consider an online stochastic game with risk-averse agents whose goal is to learn optimal decisions that minimize the risk of incurring significantly high costs. Specifically, we use the Conditional Value at Risk (CVaR) as a risk measure…

Machine Learning · Computer Science 2022-06-17 Zifan Wang , Yi Shen , Michael M. Zavlanos

We present extensive evidence that ``risk premium'' is strongly correlated with tail-risk skewness but very little with volatility. We introduce a new, intuitive definition of skewness and elicit an approximately linear relation between the…

General Finance · Quantitative Finance 2015-11-02 Y. Lempérière , C. Deremble , T. T. Nguyen , P. Seager , M. Potters , J. P. Bouchaud