Related papers: Exploratory mean-variance portfolio selection with…
Reinforcement Learning with Verifiable Rewards (RLVR) for Multimodal Large Language Models (MLLMs) is highly dependent on high-quality labeled data, which is often scarce and prone to substantial annotation noise in real-world scenarios.…
Market makers continuously set bid and ask quotes for the stocks they have under consideration. Hence they face a complex optimization problem in which their return, based on the bid-ask spread they quote and the frequency at which they…
We derive a Cram\'er-Rao lower bound for the variance of Floquet multiplier estimates that have been constructed from stable limit cycles perturbed by noise. To do so, we consider perturbed periodic orbits in the plane. We use a periodic…
The applicability of reinforcement learning (RL) algorithms in real-world domains often requires adherence to safety constraints, a need difficult to address given the asymptotic nature of the classic RL optimization objective. In contrast…
We propose a comprehensive framework for policy gradient methods tailored to continuous time reinforcement learning. This is based on the connection between stochastic control problems and randomised problems, enabling applications across…
A compact version of the variation evolving method (VEM) is developed in the primal variable space for optimal control computation. Following the idea that originates from the Lyapunov continuous-time dynamics stability theory in the…
The optimization of large portfolios displays an inherent instability to estimation error. This poses a fundamental problem, because solutions that are not stable under sample fluctuations may look optimal for a given sample, but are, in…
We consider the problem of learning a control policy that is robust against the parameter mismatches between the training environment and testing environment. We formulate this as a distributionally robust reinforcement learning (DR-RL)…
We study Markowitz's mean-variance portfolio selection problem in a continuous-time Black-Scholes market with different borrowing and saving rates. The associated Hamilton-Jacobi-Bellman equation is fully nonlinear. Using a delicate partial…
This paper introduces a multi-level (m-lev) mechanism into Evolution Strategies (ESs) in order to address a class of global optimization problems that could benefit from fine discretization of their decision variables. Such problems arise…
We investigate local optimality conditions of first and second order for integer optimal control problems with total variation regularization via a finite-dimensional switching point problem. We show the equivalence of local optimality for…
The stochastic variational inequality problem (SVIP) is an equilibrium model that includes random variables and has been widely applied in various fields such as economics and engineering. Expected residual minimization (ERM) is an…
Many potential applications of reinforcement learning (RL) require guarantees that the agent will perform well in the face of disturbances to the dynamics or reward function. In this paper, we prove theoretically that maximum entropy…
This paper first describes a class of uncertain stochastic control systems with Markovian switching, and derives an It\^o-Liu formula for Markov-modulated processes. And we characterize an optimal control law, which satisfies the…
We consider a time-consistent mean-variance portfolio selection problem of an insurer and allow for the incorporation of basis (mortality) risk. The optimal solution is identified with a Nash subgame perfect equilibrium. We characterize an…
In this note, we study a class of indefinite stochastic McKean-Vlasov linear-quadratic (LQ in short) control problem under the control taking nonnegative values. In contrast to the conventional issue, both the classical dynamic programming…
Value function based reinforcement learning (RL) algorithms, for example, $Q$-learning, learn optimal policies from datasets of actions, rewards, and state transitions. However, when the underlying state transition dynamics are stochastic…
We consider a regularized expected reward optimization problem in the non-oblivious setting that covers many existing problems in reinforcement learning (RL). In order to solve such an optimization problem, we apply and analyze the…
Gradient-regularized value learning methods improve sample efficiency by leveraging learned models of transition dynamics and rewards to estimate return gradients. However, existing approaches, such as MAGE, struggle in stochastic or noisy…
This paper examines a continuous time intertemporal consumption and portfolio choice problem with a stochastic differential utility preference of Epstein-Zin type for a robust investor, who worries about model misspecification and seeks…