相关论文: Intertemporal Hedging Demand under Epstein-Zin Pre…
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an important paradigm for unlocking reasoning capabilities in large language models, exemplified by the success of OpenAI o1 and DeepSeek-R1. Currently, Group Relative…
This thesis applies entropy as a model independent measure to address three research questions concerning financial time series. In the first study we apply transfer entropy to drawdowns and drawups in foreign exchange rates, to study their…
This paper studies the optimal risk-averse timing to sell a risky asset. The investor's risk preference is described by the exponential, power, or log utility. Two stochastic models are considered for the asset price -- the geometric…
This paper considers consumption and portfolio optimization problems with recursive preferences in both infinite and finite time regions. Specially, the financial market consists of a risk-free asset and a risky asset that follows a general…
We present a reinforcement-learning (RL) framework for dynamic hedging of equity index option exposures under realistic transaction costs and position limits. We hedge a normalized option-implied equity exposure (one unit of underlying…
This paper considers a newly delayed reinsurance and investment optimization problem incorporating random risk aversion, in which an insurer pursues maximization of the expected certainty equivalent of her/his terminal wealth and the…
The Merton investment-consumption problem is fundamental, both in the field of finance, and in stochastic control. An important extension of the problem adds transaction costs, which is highly relevant from a financial perspective but also…
Risk aversion is a key element of utility maximizing hedge strategies; however, it has typically been assigned an arbitrary value in the literature. This paper instead applies a GARCH-in-Mean (GARCH-M) model to estimate a time-varying…
This paper investigates the deep hedging framework, based on reinforcement learning (RL), for the dynamic hedging of swaptions, contrasting its performance with traditional sensitivity-based rho-hedging. We design agents under three…
In a reinforcement learning (RL) framework, we study the exploratory version of the continuous time expected utility (EU) maximization problem with a portfolio constraint that includes widely-used financial regulations such as short-selling…
Reinforcement learning from verifiable rewards has significantly advanced the reasoning capabilities of large language models. However, Group Relative Policy Optimization (GRPO) typically assigns a uniform, sequence-level advantage to all…
This thesis develops theoretical frameworks and algorithms that advance constrained reinforcement learning (RL) across control, preference learning, and alignment of large language models. The first contribution addresses constrained Markov…
We study continuous-time mean--variance portfolio selection in markets where stock prices are diffusion processes driven by observable factors that are also diffusion processes, yet the coefficients of these processes are unknown. Based on…
Preference-based Reinforcement Learning (PbRL) replaces reward values in traditional reinforcement learning by preferences to better elicit human opinion on the target objective, especially when numerical reward values are hard to design or…
We propose a hedging approach for general contingent claims when liquidity is a concern and trading is subject to transaction cost. Multiple assets with different liquidity levels are available for hedging. Our risk criterion targets a…
We introduce a novel approach to solving the optimal portfolio choice problem under Epstein-Zin utility with a time-varying consumption constraint, where analytical expressions for the value function and the dual value function are not…
A leveraged exchange traded fund (LETF) is an exchange traded fund that uses financial derivatives to amplify the price changes of a basket of goods. In this paper, we consider the robust hedging of European options on a LETF, finding…
Reinforcement Learning from Human Feedback (RLHF) has recently surged in popularity, particularly for aligning large language models and other AI systems with human intentions. At its core, RLHF can be viewed as a specialized instance of…
We study the finite mutually exclusive outcome version of risk-constrained Kelly optimization with explicit state prices. The market has outcome probabilities $p_i>0$, state prices $q_i>0$, terminal wealths $W_i=c+x_i/q_i$, and a…
This paper investigates optimal investment and pension policies in a Pay-As-You-Go (PAYG) system supplemented by a buffer fund used as an intergenerational risk-sharing mechanism. The social planner's preference criterion is represented by…