Related papers: Exploratory Utility Maximization Problem with Tsal…
This paper addresses the problem of learning optimal control policies for systems with uncertain dynamics and high-level control objectives specified as Linear Temporal Logic (LTL) formulas. Uncertainty is considered in the workspace…
In this article we provide initial findings regarding the problem of solving likelihood equations by means of a maximum entropy approach. Unlike standard procedures that require equating at zero the score function of the maximum-likelihood…
We propose to solve large scale Markowitz mean-variance (MV) portfolio allocation problem using reinforcement learning (RL). By adopting the recently developed continuous-time exploratory control framework, we formulate the exploratory MV…
In this paper we study the problem of maximizing expected utility from the terminal wealth with proportional transaction costs and random endowment. In the context of the existence of consistent price systems, we consider the duality…
Many real-world problems can be reduced to combinatorial optimization on a graph, where the subset or ordering of vertices that maximize some objective function must be found. With such tasks often NP-hard and analytically intractable,…
Known as two cornerstones of problem solving by search, exploitation and exploration are extensively discussed for implementation and application of evolutionary algorithms (EAs). However, only a few researches focus on evaluation and…
Exploration is critical to a reinforcement learning agent's performance in its given environment. Prior exploration methods are often based on using heuristic auxiliary predictions to guide policy behavior, lacking a mathematically-grounded…
While learning in an unknown Markov Decision Process (MDP), an agent should trade off exploration to discover new information about the MDP, and exploitation of the current knowledge to maximize the reward. Although the agent will…
The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. Existing algorithms for finite state and action discounted RL problems address this by assuming sufficient…
In this paper, we develop a dynamic exploration/ exploitation (exr/exp) strategy for contextual recommender systems (CRS). Specifically, our methods can adaptively balance the two aspects of exr/exp by automatically learning the optimal…
We study power utility maximization for exponential L\'evy models with portfolio constraints, where utility is obtained from consumption and/or terminal wealth. For convex constraints, an explicit solution in terms of the L\'evy triplet is…
In this paper, we investigate the robust optimal reinsurance,investment,and internal surplus distribution (i.e., consumption) problem for an insurer with Epstein-Zin recursive preferences in an incomplete market. It is assumed that the…
We consider the problem of maximising expected utility from terminal wealth in a semimartingale setting, where the semimartingale is written as a sum of a time-changed Brownian motion and a finite variation process. To solve this problem,…
Minimization problems with respect to a one-parameter family of generalized relative entropies are studied. These relative entropies, which we term relative $\alpha$-entropies (denoted $\mathscr{I}_{\alpha}$), arise as redundancies under…
The explore{exploit dilemma is one of the central challenges in Reinforcement Learning (RL). Bayesian RL solves the dilemma by providing the agent with information in the form of a prior distribution over environments; however, full…
Earlier studies have shown that stock market distributions can be well described by distributions derived from Tsallis entropy, which is a generalization of Shannon entropy to non-extensive systems. In this paper, Tsallis relative entropy…
The Lagrangian technique of Niven (2004, Physica A, 334(3-4): 444) is used to determine the constrained forms of the Tsallis entropy function - i.e. Lagrangian functions in which the probabilities of each state are independent - for each…
Safety exploration can be regarded as a constrained Markov decision problem where the expected long-term cost is constrained. Previous off-policy algorithms convert the constrained optimization problem into the corresponding unconstrained…
An important facet of reinforcement learning (RL) has to do with how the agent goes about exploring the environment. Traditional exploration strategies typically focus on efficiency and ignore safety. However, for practical applications,…
We solve a min-max problem in a robust exploratory mean-variance problem with drift uncertainty in this paper. It is verified that robust investors choose the Sharpe ratio with minimal $L^2$ norm in an admissible set. A reinforcement…