English
Related papers

Related papers: Exploratory mean-variance portfolio selection with…

200 papers

In this paper, we consider the portfolio optimization problem in a financial market where the underlying stochastic volatility model is driven by n-dimensional Brownian motions. At first, we derive a Hamilton-Jacobi-Bellman equation…

Mathematical Finance · Quantitative Finance 2024-12-20 Minglian Lin , Indranil SenGupta

We consider the portfolio optimisation problem where the terminal function is an S-shaped utility applied at the difference between the wealth and a random benchmark process. We develop several numerical methods for solving the problem…

Computational Finance · Quantitative Finance 2024-10-10 Ashley Davey , Harry Zheng

Designing model-free algorithms for distributionally robust reinforcement learning (DRRL) poses fundamental challenges. The robust Bellman operator is nonlinear in the transition kernel, which makes one-sample Bellman updates biased, while…

Machine Learning · Computer Science 2026-05-12 Shengbo Wang , Zexi Zhang

Reinforcement learning (RL) with continuous time and state/action spaces is often data-intensive and brittle under nuisance variability and shift, motivating methods that exploit value-preserving structures to stabilize and improve…

Machine Learning · Computer Science 2026-05-08 Zuyuan Zhang , Fei Xu Yu , Tian Lan

Recent advances in reinforcement learning (RL) have improved the reasoning capabilities of large language models (LLMs) and vision-language models (VLMs). However, the widely used Group Relative Policy Optimization (GRPO) consistently…

Artificial Intelligence · Computer Science 2026-04-20 Chen Wang , Lai Wei , Yanzhi Zhang , Chenyang Shao , Zedong Dan , Weiran Huang , Ge Lan , Yue Wang

Exploration is a crucial and distinctive aspect of reinforcement learning (RL) that remains a fundamental open problem. Several methods have been proposed to tackle this challenge. Commonly used methods inject random noise directly into the…

Machine Learning · Computer Science 2024-11-06 Sebastian Griesbach , Carlo D'Eramo

We investigate exploratory randomization for an extended linear-exponential-quadratic-Gaussian (LEQG) control problem in discrete time. This extended control problem is related to the structure of risk-sensitive investment management…

Optimization and Control · Mathematics 2025-09-22 Sebastien Lleo , Wolfgang Runggaldier

In this paper, both dynamic mean-variance portfolio selection problems and dynamic variance hedging problems are discussed under non-Markovian framework. Explicit closed-loop equilibrium strategies of these problems are respectively…

Optimization and Control · Mathematics 2018-02-06 Tianxiao Wang

Reinforcement learning (RL) is a powerful framework for decision-making in uncertain environments, but it often requires large amounts of data to learn an optimal policy. We address this challenge by incorporating prior model knowledge to…

Machine Learning · Computer Science 2026-01-29 J. S. van Hulst , W. P. M. H. Heemels , D. J. Antunes

In recent years, reinforcement learning (RL) systems with general goals beyond a cumulative sum of rewards have gained traction, such as in constrained problems, exploration, and acting upon prior experiences. In this paper, we consider…

Machine Learning · Computer Science 2020-07-07 Junyu Zhang , Alec Koppel , Amrit Singh Bedi , Csaba Szepesvari , Mengdi Wang

In this article, two methods for solving mean-field type optimal control problems are proposed and investigated. The two methods are iterative methods: at each iteration, a Hamilton-Jacobi-Bellman equation is solved, for a terminal…

Optimization and Control · Mathematics 2017-03-30 Laurent Pfeiffer

We present a computational motivation for restricted maximum likelihood (REML) estimation in linear mixed models using an expectation--maximization (EM) algorithm. At each iteration, maximum likelihood (ML) and REML solve the same…

Computation · Statistics 2026-02-11 Andrew T. Karl

In the present work, we present numerical results for an iterative method for solving an optimal control problem with inequality contraints. The method is based on generalized Bregman distances. Under a combination of a source condition and…

Optimization and Control · Mathematics 2016-06-07 Frank Pörner

Reinforcement learning (RL) algorithms typically optimize the expected cumulative reward, i.e., the expected value of the sum of scalar rewards an agent receives over the course of a trajectory. The expected value averages the performance…

Machine Learning · Computer Science 2025-09-01 Xinyi Sheng , Dominik Baumann

The exploration \& exploitation dilemma poses significant challenges in reinforcement learning (RL). Recently, curiosity-based exploration methods achieved great success in tackling hard-exploration problems. However, they necessitate…

Machine Learning · Computer Science 2024-12-06 Yiran Wang , Chenshu Liu , Yunfan Li , Sanae Amani , Bolei Zhou , Lin F. Yang

Reinforcement learning algorithms solve sequential decision-making problems in probabilistic environments by optimizing for long-term reward. The desire to use reinforcement learning in safety-critical settings inspires a recent line of…

Artificial Intelligence · Computer Science 2020-12-17 Koundinya Vajjha , Avraham Shinnar , Vasily Pestun , Barry Trager , Nathan Fulton

We study the problem of estimating Dynamic Discrete Choice (DDC) models, also known as offline Maximum Entropy-Regularized Inverse Reinforcement Learning (offline MaxEnt-IRL) in machine learning. The objective is to recover reward or $Q^*$…

Machine Learning · Computer Science 2026-05-06 Enoch H. Kang , Hema Yoganarasimhan , Lalit Jain

The monotone mean-variance (MMV) preference proposed by Maccheroni, et al. (Math. Finance 19(3): 487-521, 2009) fails to differentiate strictly dominant payoffs, which may cause inconsistency in portfolio decision-making. This paper…

Mathematical Finance · Quantitative Finance 2026-04-03 Yike Wang , Yusha Chen , Jingzhen Liu , Zhenyu Cui

Optimized certainty equivalents (OCEs) is a family of risk measures widely used by both practitioners and academics. This is mostly due to its tractability and the fact that it encompasses important examples, including entropic risk…

Optimization and Control · Mathematics 2022-06-07 Julio Backhoff Veraguas , A. Max Reppen , Ludovic Tangpi

Balancing exploration and exploitation remains a central challenge in reinforcement learning with verifiable rewards (RLVR) for large language models (LLMs). Current RLVR methods often overemphasize exploitation, leading to entropy…

Computation and Language · Computer Science 2026-04-14 Liang Chen , Xueting Han , Qizhou Wang , Bo Han , Jing Bai , Hinrich Schutze , Kam-Fai Wong
‹ Prev 1 4 5 6 7 8 10 Next ›