English
Related papers

Related papers: Breaking the Dimensional Barrier for Constrained D…

200 papers

We introduce the Pontryagin-Guided Direct Policy Optimization (PG-DPO) framework for high-dimensional continuous-time portfolio choice. Our approach combines Pontryagin's Maximum Principle (PMP) with backpropagation through time (BPTT) to…

Portfolio Management · Quantitative Finance 2025-09-12 Jeonggyu Huh , Jaegi Jeon , Hyeng Keun Koo , Byung Hwa Lim

We present a Pontryagin-Guided Direct Policy Optimization (PG-DPO) framework for Merton's portfolio problem, unifying modern neural-network-based policy parameterization with the adjoint viewpoint from Pontryagin's maximum principle (PMP).…

Optimization and Control · Mathematics 2025-01-14 Jeonggyu Huh , Jaegi Jeon

We study the problem of optimal portfolio selection under stochastic volatility within a continuous time reinforcement learning framework with portfolio constraints. Exploration is modeled through entropy-regularized relaxed controls, where…

Mathematical Finance · Quantitative Finance 2026-04-27 Thai Nguyen , Pertiny Nkuize

Most value-based and actor--critic reinforcement learning methods rely on Bellman-style recursions, yet these recursions collapse under non-exponential discounting common in human preferences and survival processes. We show the breakdown is…

Machine Learning · Computer Science 2026-05-21 Hojin Ko , Jeonggyu Huh

We study continuous-time CRRA portfolio choice in diffusion markets with estimated and hence uncertain coefficients. Nature draws a latent parameter $\theta \sim q$ at time $0$ and keeps it fixed; the investor never observes $\theta$ and…

Computational Finance · Quantitative Finance 2026-02-27 Jeonggyu Huh , Hyeng Keun Koo

A Hamiltonian algorithm, both theoretical and numerical, to obtain the reduced equations implementing Pontryagine's Maximum Principle for singular linear-quadratic optimal control problems is presented. This algorithm is inspired on the…

Optimization and Control · Mathematics 2012-04-13 M. Delgado-Tellez , A. Ibort

Policy gradient methods are widely used in reinforcement learning. Yet, the nonconvexity of policy optimization poses significant challenges in understanding the global convergence of policy gradient methods. For a class of finite-horizon…

Optimization and Control · Mathematics 2026-03-10 Xin Chen , Yifan Hu , Minda Zhao

In this article we derive a Pontryagin maximum principle (PMP) for discrete-time optimal control problems on matrix Lie groups. The PMP provides first order necessary conditions for optimality; these necessary conditions typically yield two…

Systems and Control · Computer Science 2018-08-07 Karmvir Singh Phogat , Debasish Chatterjee , Ravi Banavar

Stochastic control problems in high dimensions are notoriously difficult to solve due to the curse of dimensionality. An alternative to traditional dynamic programming is Pontryagin's Maximum Principle (PMP), which recasts the problem as a…

Machine Learning · Computer Science 2025-07-03 Qian Qi

Constrained decision-making is essential for designing safe policies in real-world control systems, yet simulated environments often fail to capture real-world adversities. We consider the problem of learning a policy that will maximize the…

Machine Learning · Computer Science 2026-02-10 Sourav Ganguly , Kishan Panaganti , Arnob Ghosh , Adam Wierman

In this article we derive a strong version of the Pontryagin Maximum Principle for general nonlinear optimal control problems on time scales in finite dimension. The final time can be fixed or not, and in the case of general boundary…

Optimization and Control · Mathematics 2013-02-15 Loïc Bourdin , Emmanuel Trélat

Exploiting our previous results on higher order controlled Lagrangians in [Nonlinear Anal. {\bf 207} (2021), 112263], we derive here an analogue of the classical first order Pontryagin Maximum Principle (PMP) for cost minimising problems…

Optimization and Control · Mathematics 2023-03-17 Franco Cardin , Cristina Giannotti , Andrea Spiro

The continuous dynamical system approach to deep learning is explored in order to devise alternative frameworks for training algorithms. Training is recast as a control problem and this allows us to formulate necessary optimality conditions…

Machine Learning · Computer Science 2018-06-05 Qianxiao Li , Long Chen , Cheng Tai , Weinan E

We propose a neural network approach that yields approximate solutions for high-dimensional optimal control problems and demonstrate its effectiveness using examples from multi-agent path finding. Our approach yields controls in a feedback…

Optimization and Control · Mathematics 2022-06-29 Derek Onken , Levon Nurbekyan , Xingjian Li , Samy Wu Fung , Stanley Osher , Lars Ruthotto

We study policy optimization in an infinite horizon, $\gamma$-discounted constrained Markov decision process (CMDP). Our objective is to return a policy that achieves large expected reward with a small constraint violation. We consider the…

Machine Learning · Computer Science 2022-04-12 Arushi Jain , Sharan Vaswani , Reza Babanezhad , Csaba Szepesvari , Doina Precup

The main objective of this paper is to develop a martingale-type solution to optimal consumption--investment choice problems ([Merton, 1969] and [Merton, 1971]) under time-varying incomplete preferences driven by externalities such as…

Mathematical Finance · Quantitative Finance 2025-01-14 Weixuan Xia

We consider approximate dynamic programming for the infinite-horizon stationary $\gamma$-discounted optimal control problem formalized by Markov Decision Processes. While in the exact case it is known that there always exists an optimal…

Optimization and Control · Mathematics 2013-04-23 Boris Lesner , Bruno Scherrer

We consider a dynamic portfolio optimization problem that incorporates predictable returns, instantaneous transaction costs, price impact, and stochastic volatility, extending the classical results of Garleanu and Pedersen (2013), which…

Computational Finance · Quantitative Finance 2025-07-24 Patrick Chan , Ronnie Sircar , Iosif Zimbidis

This paper investigates a continuous-time portfolio optimization problem with the following features: (i) a no-short selling constraint; (ii) a leverage constraint, that is, an upper limit for the sum of portfolio weights; and (iii) a…

Portfolio Management · Quantitative Finance 2022-03-08 Masashi Ieda

We study episodic reinforcement learning (RL) in non-stationary linear kernel Markov decision processes (MDPs). In this setting, both the reward function and the transition kernel are linear with respect to the given feature maps and are…

Machine Learning · Computer Science 2024-12-24 Han Zhong , Zhongren Chen , Zhuoran Yang , Zhaoran Wang , Csaba Szepesvári
‹ Prev 1 2 3 10 Next ›