English
Related papers

Related papers: Breaking the Dimensional Barrier for Constrained D…

200 papers

A standard objective in partially-observable Markov decision processes (POMDPs) is to find a policy that maximizes the expected discounted-sum payoff. However, such policies may still permit unlikely but highly undesirable outcomes, which…

Artificial Intelligence · Computer Science 2017-01-31 Krishnendu Chatterjee , Petr Novotný , Guillermo A. Pérez , Jean-François Raskin , Đorđe Žikelić

We propose a framework, called neural-progressive hedging (NP), that leverages stochastic programming during the online phase of executing a reinforcement learning (RL) policy. The goal is to ensure feasibility with respect to constraints…

Machine Learning · Computer Science 2022-03-01 Supriyo Ghosh , Laura Wynter , Shiau Hong Lim , Duc Thien Nguyen

We study the time-optimal robust control of a two-level quantum system subjected to field inhomogeneities. We apply the Pontryagin Maximum Principle and we introduce a reduced space onto which the optimal dynamics is projected down. This…

Quantum Physics · Physics 2025-09-03 O. Fresse-Colson , S. Guérin , Xi Chen , D. Sugny

This paper examines a continuous time intertemporal consumption and portfolio choice problem with a stochastic differential utility preference of Epstein-Zin type for a robust investor, who worries about model misspecification and seeks…

Optimization and Control · Mathematics 2021-03-09 Jiangyan Pu , Qi Zhang

This paper develops a robust fixed time optimization framework for constrained problems that guarantees exact constraint satisfaction and convergence to KKT points within fixed time , independent of initial conditions. The approach treats…

Optimization and Control · Mathematics 2026-05-27 Baby Diana , Priyanka Singh , Shyam Kamal , Sandip Ghosh , Bijnan Bandyopadhyay

Many control problems in environments that can be modeled as Markov decision processes (MDPs) concern infinite-time horizon specifications. The classical aim in this context is to compute a control policy that maximizes the probability of…

Systems and Control · Computer Science 2017-05-03 Ruediger Ehlers , Salar Moarref , Ufuk Topcu

Based on Pontryagin Maximum Principle (PMP), this paper establishes a generalized PMP aiming at control system with with extra input/output terms. The paper details the adaptive target and gives a proof of the generalized theorem.…

Optimization and Control · Mathematics 2016-01-01 Yuanzun Zhao

We consider the infinite-horizon discounted optimal control problem formalized by Markov Decision Processes. We focus on several approximate variations of the Policy Iteration algorithm: Approximate Policy Iteration, Conservative Policy…

Artificial Intelligence · Computer Science 2014-05-13 Bruno Scherrer

We introduce \texttt{OPO-CMDP}, the first policy optimization algorithm for stochastic Contextual Markov Decision Process (CMDPs) under general offline function approximation. Our approach achieves a high probability regret bound of…

Machine Learning · Computer Science 2026-02-17 Orin Levy , Aviv Rosenberg , Alon Cohen , Yishay Mansour

We consider a portfolio optimization problem in a defaultable market with finitely-many economical regimes, where the investor can dynamically allocate her wealth among a defaultable bond, a stock, and a money market account. The market…

Portfolio Management · Quantitative Finance 2011-09-07 Agostino Capponi , Jose E. Figueroa-Lopez

Proximal policy optimization(PPO) has been proposed as a first-order optimization method for reinforcement learning. We should notice that an exterior penalty method is used in it. Often, the minimizers of the exterior penalty functions…

Machine Learning · Computer Science 2018-12-18 Cheng Zeng , Hongming Zhang

The proximal policy optimization (PPO) algorithm stands as one of the most prosperous methods in the field of reinforcement learning (RL). Despite its success, the theoretical understanding of PPO remains deficient. Specifically, it is…

Machine Learning · Computer Science 2023-06-09 Han Zhong , Tong Zhang

Policy Iteration (PI) is a widely used family of algorithms to compute optimal policies for Markov Decision Problems (MDPs). We derive upper bounds on the running time of PI on Deterministic MDPs (DMDPs): the class of MDPs in which every…

Discrete Mathematics · Computer Science 2023-10-10 Ritesh Goenka , Eashan Gupta , Sushil Khyalia , Pratyush Agarwal , Mulinti Shaik Wajid , Shivaram Kalyanakrishnan

Proximal Policy Optimization (PPO) has become the predominant algorithm for on-policy reinforcement learning due to its scalability and empirical robustness across domains. However, there is a significant disconnect between the underlying…

In this paper, we derive a version of the Pontryagin maximum principle for general finite-dimensional nonlinear optimal sampled-data control problems. Our framework is actually much more general, and we treat optimal control problems for…

Optimization and Control · Mathematics 2015-12-09 Loïc Bourdin , Emmanuel Trélat

We consider nonsmooth optimal control problems subject to a linear elliptic partial differential equation with homogeneous Dirichlet boundary conditions. It is well-known that local solutions satisfy the celebrated Pontryagin maximum…

Optimization and Control · Mathematics 2024-06-28 Daniel Wachsmuth

I study intertemporal hedging demand in a continuous-time multi-asset long-run risk (LRR) model under Epstein--Zin (EZ) recursive preferences. The investor trades a risk-free asset and several risky assets whose drifts and volatilities…

Systems and Control · Electrical Eng. & Systems 2025-12-18 Wonchan Cho

Many processes, such as discrete event systems in engineering or population dynamics in biology, evolve in discrete space and continuous time. We consider the problem of optimal decision making in such discrete state and action space…

Machine Learning · Computer Science 2020-10-27 Bastian Alt , Matthias Schultheis , Heinz Koeppl

In the optimization of dynamic systems, the variables typically have constraints. Such problems can be modeled as a Constrained Markov Decision Process (CMDP). This paper considers the peak Constrained Markov Decision Process (PCMDP), where…

Optimization and Control · Mathematics 2022-06-15 Qinbo Bai , Vaneet Aggarwal , Ather Gattami

While the techniques in optimal control theory are often model-based, the policy optimization (PO) approach directly optimizes the performance metric of interest. Even though it has been an essential approach for reinforcement learning…

Optimization and Control · Mathematics 2022-11-23 Feiran Zhao , Keyou You , Tamer Başar
‹ Prev 1 4 5 6 7 8 10 Next ›