中文
相关论文

相关论文: Intertemporal Hedging Demand under Epstein-Zin Pre…

200 篇论文

Stochastic and soft optimal policies resulting from entropy-regularized Markov decision processes (ER-MDP) are desirable for exploration and imitation learning applications. Motivated by the fact that such policies are sensitive with…

机器学习 · 计算机科学 2022-01-03 Tien Mai , Patrick Jaillet

In finance, sequential decision problems are often faced, for which reinforcement learning (RL) emerges as a promising tool for optimisation without the need of analytical tractability. However, the objective of classical RL is the expected…

计算金融 · 定量金融 2026-02-13 Federico Cacciamani , Roberto Daluiso , Marco Pinciroli , Michele Trapletti , Edoardo Vittori

Bootstrapping large language models (LLMs) through preference-based policy optimization offers a promising direction for aligning model behavior with human preferences without relying on extensive manual annotations. In this work, we…

人工智能 · 计算机科学 2025-12-25 Chen Jia

Numerous empirical proofs indicate the adequacy of the time discrete auto-regressive stochastic volatility models introduced by Taylor in the description of the log-returns of financial assets. The pricing and hedging of contingent products…

证券定价 · 定量金融 2011-10-31 Joan del Castillo , Juan-Pablo Ortega

We study episodic reinforcement learning (RL) in non-stationary linear kernel Markov decision processes (MDPs). In this setting, both the reward function and the transition kernel are linear with respect to the given feature maps and are…

机器学习 · 计算机科学 2024-12-24 Han Zhong , Zhongren Chen , Zhuoran Yang , Zhaoran Wang , Csaba Szepesvári

The objectives of option hedging/trading extend beyond mere protection against downside risks, with a desire to seek gains also driving agent's strategies. In this study, we showcase the potential of robust risk-aware reinforcement learning…

计算金融 · 定量金融 2023-12-27 David Wu , Sebastian Jaimungal

Deep Reinforcement Learning (DRL) is a powerful tool used for addressing complex challenges in mobile networks. This paper investigates the application of two DRL models, on-policy and off-policy, in the field of resource allocation for…

网络与互联网体系结构 · 计算机科学 2024-12-04 Manal Mehdaoui , Amine Abouaomar

Offline Preference-based Reinforcement Learning (PbRL) learns rewards and policies aligned with human preferences without the need for extensive reward engineering and direct interaction with human annotators. However, ensuring safety…

人工智能 · 计算机科学 2025-12-24 Ze Gong , Pradeep Varakantham , Akshat Kumar

Research in quantitative finance has demonstrated that reinforcement learning (RL) methods have delivered promising outcomes in the context of hedging financial portfolios. For example, hedging a portfolio of European options using RL…

计算工程、金融与科学 · 计算机科学 2024-07-16 Anil Sharma , Freeman Chen , Jaesun Noh , Julio DeJesus , Mario Schlener

A rational behavior of a consumer is analyzed when the user participates in a Peak Time Rebate (PTR) mechanism, which is a demand response (DR) incentive program based on a baseline. A multi-stage stochastic programming is proposed from the…

系统与控制 · 计算机科学 2018-02-23 José Vuelvas , Fredy Ruiz

We study the optimal scheduling problem for a Markovian multiclass queueing network with abandonment in the Halfin--Whitt regime, under the long run average (ergodic) risk sensitive cost criterion. The objective is to prove asymptotic…

概率论 · 数学 2024-10-23 Sumith Reddy Anugu , Guodong Pang

A macroeconomic model based on the economic variables (i) assets, (ii) leverage (defined as debt over asset) and (iii) trust (defined as the maximum sustainable leverage) is proposed to investigate the role of credit in the dynamics of…

经济学 · 定量金融 2016-08-24 Jeroen Rozendaal , Yannick Malevergne , Didier Sornette

This paper proposes and analyzes two new policy learning methods: regularized policy gradient (RPG) and iterative policy optimization (IPO), for a class of discounted linear-quadratic control (LQC) problems over an infinite time horizon…

最优化与控制 · 数学 2025-10-08 Xin Guo , Xinyu Li , Renyuan Xu

This work extends a previous work in regime detection, which allowed trading positions to be profitably adjusted when a new regime was detected, to ex ante prediction of regimes, leading to substantial performance improvements over the…

风险管理 · 定量金融 2023-10-10 Piotr Pomorski , Denise Gorse

Hedging a portfolio containing autocallable notes presents unique challenges due to the complex risk profile of these financial instruments. In addition to hedging, pricing these notes, particularly when multiple underlying assets are…

计算工程、金融与科学 · 计算机科学 2024-11-05 Anil Sharma , Freeman Chen , Jaesun Noh , Julio DeJesus , Mario Schlener

We investigate the Distributionally Robust Regret-Optimal (DR-RO) control of discrete-time linear dynamical systems with quadratic cost over an infinite horizon. Regret is the difference in cost obtained by a causal controller and a…

系统与控制 · 电气工程与系统科学 2024-01-01 Taylan Kargin , Joudi Hajar , Vikrant Malik , Babak Hassibi

Prepayment risk embedded in fixed-rate mortgages forms a significant fraction of a financial institution's exposure, and it receives particular attention because of the magnitude of the underlying market. The embedded prepayment option…

计算金融 · 定量金融 2024-10-29 Leonardo Perotti , Lech A. Grzelak , Cornelis W. Oosterlee

This study develops a regime-aware portfolio allocation framework that integrates Markov switching models with Reinforcement Learning (RL) to dynamically allocate across equities (SPY), long-term Treasuries (TLT), and gold (GLD). Using…

投资组合管理 · 定量金融 2026-05-28 Ajay Kumar Verma , Nunik Srikandi Putri , Neo Paul Lesupi

Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates significant potential in enhancing the reasoning capabilities of Large Language Models (LLMs). However, existing RLVR methods are often constrained by issues such as…

人工智能 · 计算机科学 2026-01-14 Jinpeng Wang , Chao Li , Ting Ye , Mengyuan Zhang , Wei Liu , Jian Luan

A class of heterogeneous agent models is investigated where investors switch trading position whenever their motivation to do so exceeds some critical threshold. These motivations can be psychological in nature or reflect behaviour…

计算物理 · 物理学 2009-11-13 R. Cross , M. Grinfeld , H. Lamba , T. Seaman