Related papers: Robust exploratory mean-variance problem with drif…
A reinforcement learning agent tries to maximize its cumulative payoff by interacting in an unknown environment. It is important for the agent to explore suboptimal actions as well as to pick actions with highest known rewards. Yet, in…
Options are generally learned by using an inaccurate environment model (or simulator), which contains uncertain model parameters. While there are several methods to learn options that are robust against the uncertainty of model parameters,…
This paper studies a robust stochastic control problem with a monotone mean-variance cost functional and random coefficients. The main technique is to find the saddle point through two backward stochastic differential equations (BSDEs) with…
Stochastic and soft optimal policies resulting from entropy-regularized Markov decision processes (ER-MDP) are desirable for exploration and imitation learning applications. Motivated by the fact that such policies are sensitive with…
Monotone mean-variance (MMV) utility is the minimal modification of the classical Markowitz utility that respects rational ordering of investment opportunities. This paper provides, for the first time, a complete characterization of optimal…
This paper solves the dynamic portfolio choice problem. Using an explicit solution with a power utility, we construct a bridge between a continuous and discrete VAR model to assess portfolio sensitivities. We find, from a well analyzed…
Extreme Value Theory (EVT) is one of the most commonly used approaches in finance for measuring the downside risk of investment portfolios, especially during financial crises. In this paper, we propose a novel approach based on EVT called…
This paper expands the notion of robust profit opportunities in financial markets to incorporate distributional uncertainty using Wasserstein distance as the ambiguity measure. Financial markets with risky and risk-free assets are…
We study continuous-time mean--variance portfolio selection in markets where stock prices are diffusion processes driven by observable factors that are also diffusion processes, yet the coefficients of these processes are unknown. Based on…
We study hedging and pricing of unattainable contingent claims in a non-Markovian regime-switching financial model. Our financial market consists of a bank account and a risky asset whose dynamics are driven by a Brownian motion and a…
In this paper we study mean-variance hedging under the G-expectation framework. Our analysis is carried out by exploiting the G-martingale representation theorem and the related probabilistic tools, in a contin- uous financial market with…
This paper proposes a new robust optimization (RO) formulation namely the RO under objective functional uncertainty (ObRO). The ObRO adopts a min-max structure where the inner problem finds the worst-case objective function in a continuous…
Our goal is to build robust optimization problems for making decisions based on complex data from the past. In robust optimization (RO) generally, the goal is to create a policy for decision-making that is robust to our uncertainty about…
Exploration of unknown, unstructured environments, such as in search and rescue, cave exploration, and planetary missions,presents significant challenges due to their unpredictable nature. This unpredictability can lead to inefficient path…
We study issues of robustness in the context of Quantitative Risk Management and Optimization. We develop a general methodology for determining whether a given risk measurement related optimization problem is robust, which we call…
In this paper, we study an optimal mean-variance investment-reinsurance problem for an insurer (she) under a Cram\'er-Lundberg model with random coefficients. At any time, the insurer can purchase reinsurance or acquire new business and…
This paper aims to put forward the concept that learning to take safe actions in unknown environments, even with probability one guarantees, can be achieved without the need for an unbounded number of exploratory trials, provided that one…
In a fixed time horizon, appropriately executing a large amount of a particular asset -- meaning a considerable portion of the volume traded within this frame -- is challenging. Especially for illiquid or even highly liquid but also highly…
We study the problem of optimal portfolio selection under stochastic volatility within a continuous time reinforcement learning framework with portfolio constraints. Exploration is modeled through entropy-regularized relaxed controls, where…
We consider the problem of learning from training data obtained in different contexts, where the underlying context distribution is unknown and is estimated empirically. We develop a robust method that takes into account the uncertainty of…