相关论文: Intertemporal Hedging Demand under Epstein-Zin Pre…
This paper presents a deep reinforcement learning (DRL) framework for dynamic portfolio optimization under market uncertainty and risk. The proposed model integrates a Sharpe ratio-based reward function with direct risk control mechanisms,…
In this paper we solve the discrete time mean-variance hedging problem when asset returns follow a multivariate autoregressive hidden Markov model. Time dependent volatility and serial dependence are well established properties of financial…
This paper investigates the integration of response time data into human preference learning frameworks for more effective reward model elicitation. While binary preference data has become fundamental in fine-tuning foundation models,…
Trust region policy optimization (TRPO) is a popular and empirically successful policy search algorithm in Reinforcement Learning (RL) in which a surrogate problem, that restricts consecutive policies to be 'close' to one another, is…
The first motivation of our paper is to explore further the idea that, in risk control problems, it may be profitable to base decisions both on the position of the underlying process Xt and on its supremum Xt := sup 0$\le$s$\le$t Xs.…
This paper studies robust forward investment and consumption preferences within a zero-volatility context. Different from previous works, we consider an incomplete financial market model due to general investment portfolio constraints. We…
We consider the classical multi-asset Merton investment problem under drift uncertainty, i.e. the asset price dynamics are given by geometric Brownian motions with constant but unknown drift coefficients. The investor assumes a prior drift…
Dynamic hedging is the practice of periodically transacting financial instruments to offset the risk caused by an investment or a liability. Dynamic hedging optimization can be framed as a sequential decision problem; thus, Reinforcement…
The main objective of this paper is to develop a martingale-type solution to optimal consumption--investment choice problems ([Merton, 1969] and [Merton, 1971]) under time-varying incomplete preferences driven by externalities such as…
In this paper, we consider a dynamic coalition portfolio selection problem, with each agent's objective given by an Epstein--Zin recursive utility. To find a Pareto optimum, the coalition's problem is formulated as an optimization problem…
In the frictionless discrete time financial market of Bouchard et al.(2015) we consider a trader who, due to regulatory requirements or internal risk management reasons, is required to hedge a claim $\xi$ in a risk-conservative way relative…
This paper investigates a novel behavioral feature of recursive preferences: aversion to risks that persist over time, or simply \textit{correlation aversion}. Greater persistence provides information about future consumption but reduces…
Traditional risk factors like beta, size/value, and momentum often lag behind market dynamics in measuring and predicting stock return volatility. Statistical models like PCA and factor analysis fail to capture hidden nonlinear…
Distributionally Robust Optimization (DRO) is a popular framework for decision-making under uncertainty, but its adversarial nature can lead to overly conservative solutions. To address this, we study ex-ante Distributionally Robust Regret…
Despite Proximal Policy Optimization (PPO) dominating policy gradient methods -- from robotic control to game AI -- its static trust region forces a brittle trade-off: aggressive clipping stifles early exploration, while late-stage updates…
Sustainable finance, which integrates environmental, social and governance (ESG) criteria on financial decisions rests on the fact that money should be used for good purposes. Thus, the financial sector is also expected to play a more…
Risk-sensitive control balances performance with resilience to unlikely events in uncertain systems. This paper introduces ergodic-risk criteria, which capture long-term cumulative risks through probabilistic limit theorems. By ensuring the…
We study portfolio selection in a complete continuous-time market where the preference is dictated by the rank-dependent utility. As such a model is inherently time inconsistent due to the underlying probability weighting, we study the…
Aggregating risks from multiple sources can be complex and demanding, and decision makers usually adopt heuristics to simplify the evaluation process. This paper axiomatizes two closed related and yet different heuristics, narrow bracketing…
This paper studies the mean-variance optimal portfolio choice of an investor pre-committed to a deterministic investment policy in continuous time in a market with mean-reversion in the risk-free rate and the equity risk-premium. In the…