Related papers: Exploratory Utility Maximization Problem with Tsal…
In this article we consider an optimization problem of expected utility maximization of continuous-time trading in a financial market. This trading is constrained by a benchmark for a utility-based shortfall risk measure. The market…
The sum-utility maximization problem is known to be important in the energy systems literature. The conventional assumption to address this problem is that the utility is concave. But for some key applications, such an assumption is not…
This paper studies the continuous-time reinforcement learning (RL) for optimal switching problems across multiple regimes. We consider a type of exploratory formulation under entropy regularization where the agent randomizes both the timing…
We investigate optimal consumption problems for a Black-Scholes market under uniform restrictions on Value-at-Risk and Expected Shortfall for logarithmic utility functions. We find the solutions in terms of a dynamic strategy in explicit…
In this paper, we study the optimal stopping problem in the so-called exploratory framework, in which the agent takes actions randomly conditioning on current state and an entropy-regularized term is added to the reward functional. Such a…
Reasoning large language models (LLMs) excel in complex tasks, which has drawn significant attention to reinforcement learning (RL) for LLMs. However, existing approaches allocate an equal number of rollouts to all questions during the RL…
It has recently been a common practice to maximize the deformed entropies through the escort averaging scheme. However, whatever averaging procedure is employed, one should recover the ordinary Shannon maximization results in the…
We provide a new algorithm for solving Risk Sensitive Partially Observable Markov Decisions Processes, when the risk is modeled by a utility function, and both the state space and the space of observations is finite. This algorithm is based…
The purpose of this note is to give the general solution of two functional equations connected to the Shannon entropy and also to the Tsallis entropy. As a result of this, we present the regular solution of these equations, as well.…
We consider the problem of learning the optimal policy for Markov decision processes with safety constraints. We formulate the problem in a reach-avoid setup. Our goal is to design online reinforcement learning algorithms that ensure safety…
In this paper, we study expected utility maximization under ratchet and drawdown constraints on consumption in a general incomplete semimartingale market using duality methods. The optimization is considered with respect to two parameters:…
We address payoff-based decentralized learning in infinite-horizon zero-sum Markov games. In this setting, each player makes decisions based solely on received rewards, without observing the opponent's strategy or actions nor sharing…
We consider a stochastic financial incomplete market where the price processes are described by a vector-valued semimartingale that is possibly nonlocally bounded. We face the classical problem of utility maximization from terminal wealth,…
This paper studies a robust utility maximization problem for intractable claims under distributional ambiguity, where the distribution of the claim cannot be inferred from market information and its dependence with tradable assets is…
In reinforcement learning (RL), an agent must explore an initially unknown environment in order to learn a desired behaviour. When RL agents are deployed in real world environments, safety is of primary concern. Constrained Markov decision…
An effective approach to exploration in reinforcement learning is to rely on an agent's uncertainty over the optimal policy, which can yield near-optimal exploration strategies in tabular settings. However, in non-tabular settings that…
Reinforcement learning (RL) is a powerful framework for decision-making in uncertain environments, but it often requires large amounts of data to learn an optimal policy. We address this challenge by incorporating prior model knowledge to…
When deploying artificial agents in real-world environments where they interact with humans, it is crucial that their behavior is aligned with the values, social norms or other requirements of that environment. However, many environments…
A drawdown constraint forces the current wealth to remain above a given function of its maximum to date. We consider the portfolio optimisation problem of maximising the long-term growth rate of the expected utility of wealth subject to a…
This paper studies the continuous time utility maximization problem on consumption with addictive habit formation in incomplete semimartingale markets. Introducing the set of auxiliary state processes and the modified dual space, we embed…