中文
相关论文

相关论文: Approximative Policy Iteration for Exit Time Feedb…

200 篇论文

This article details a general numerical framework to approximate so-lutions to linear programs related to optimal transport. The general idea is to introduce an entropic regularization of the initial linear program. This regularized…

数值分析 · 数学 2014-12-17 Jean-David Benamou , Guillaume Carlier , Marco Cuturi , Luca Nenna , Gabriel Peyré

In this paper, a convex optimization-based method is proposed for numerically solving dynamic programs in continuous state and action spaces. The key idea is to approximate the output of the Bellman operator at a particular state by the…

最优化与控制 · 数学 2020-10-23 Insoon Yang

We present a stochastic predictive controller for discrete time linear time invariant systems under incomplete state information. Our approach is based on a suitable choice of control policies, stability constraints, and employment of a…

最优化与控制 · 数学 2018-02-27 Prabhat Kumar Mishra , Debasish Chatterjee , Daniel E. Quevedo

In this paper, we introduce a model-based deep-learning approach to solve finite-horizon continuous-time stochastic control problems with jumps. We iteratively train two neural networks: one to represent the optimal policy and the other to…

机器学习 · 计算机科学 2026-01-16 Patrick Cheridito , Jean-Loup Dupret , Donatien Hainaut

Trajectory optimization is a fundamental stochastic optimal control problem. This paper deals with a trajectory optimization approach for dynamical systems subject to measurement noise that can be fitted into linear time-varying stochastic…

系统与控制 · 电气工程与系统科学 2021-08-24 Prakash Mallick , Zhiyong Chen

This paper establishes a rigorous connection between regularized discrete-time reinforcement learning (RL) and continuous-time stochastic optimal control. Specifically, classical RL algorithms are typically solving a regularized…

最优化与控制 · 数学 2026-04-24 Huyên Pham , Yuming Paul Zhang , Yuhua Zhu

We consider the stochastic optimal control problem of McKean-Vlasov stochastic differential equation where the coefficients may depend upon the joint law of the state and control. By using feedback controls, we reformulate the problem into…

概率论 · 数学 2017-03-09 Huyên Pham , Xiaoli Wei

We apply the Stochastic Perron method, created by Bayraktar and S\^irbu, to a stochastic exit time control problem. Our main assumption is the validity of the Strong Comparison Result for the related Hamilton-Jacobi-Bellman (HJB) equation.…

最优化与控制 · 数学 2013-11-01 Dmitry B. Rokhlin

We propose a principled kernel-based policy iteration algorithm to solve the continuous-state Markov Decision Processes (MDPs). In contrast to most decision-theoretic planning frameworks, which assume fully known state transition models, we…

机器人学 · 计算机科学 2020-06-04 Junhong Xu , Kai Yin , Lantao Liu

We consider infinite horizon dynamic programming problems, where the control at each stage consists of several distinct decisions, each one made by one of several agents. In an earlier work we introduced a policy iteration algorithm, where…

最优化与控制 · 数学 2020-05-05 Dimitri Bertsekas

We consider the optimal regulation problem for nonlinear control-affine dynamical systems. Whereas the linear-quadratic regulator (LQR) considers optimal control of a linear system with quadratic cost function, we study polynomial systems…

最优化与控制 · 数学 2024-10-30 Nicholas A. Corbin , Boris Kramer

The objective of this work is to study continuous-time Markov decision processes on a general Borel state space with both impulsive and continuous controls for the infinite-time horizon discounted cost. The continuous-time controlled…

最优化与控制 · 数学 2019-08-17 François Dufour , Alexei Piunovskiy

Following the recent resurgence in establishing linear control theoretic benchmarks for reinforcement leaning (RL)-based policy optimization (PO) for complex dynamical systems with continuous state and action spaces, an optimal control…

系统与控制 · 电气工程与系统科学 2023-06-30 Leilei Cui , Lekan Molu

We present a unified framework for learning continuous control policies using backpropagation. It supports stochastic control by treating stochasticity in the Bellman equation as a deterministic function of exogenous noise. The product is a…

机器学习 · 计算机科学 2015-11-02 Nicolas Heess , Greg Wayne , David Silver , Timothy Lillicrap , Yuval Tassa , Tom Erez

We revisit the linear programming approach to deterministic, continuous time, infinite horizon discounted optimal control problems. In the first part, we relax the original problem to an infinite-dimensional linear program over a measure…

最优化与控制 · 数学 2017-06-08 Angeliki Kamoutsi , Tobias Sutter , Peyman Mohajerin Esfahani , John Lygeros

In this paper, we propose a Transformer-based framework for approximating solutions to infinite-dimensional optimization problems: calculus of variations problems and optimal control problems. Our approach leverages offline training on data…

最优化与控制 · 数学 2025-11-20 Gage MacLin , Venanzio Cichella , Andrew Patterson , Irene Gregory

The theory of optimal control on positive cones has recently identified several new problem classes where the Bellman equation can be solved explicitly, in analogy with classical linear quadratic control. In this paper, the idea is extended…

最优化与控制 · 数学 2025-12-01 Anders Rantzer

This paper is concerned with the open-loop time-consistent solution of time-inconsistent mean-field stochastic linear-quadratic optimal control. Different from standard stochastic linear-quadratic problems, both the system matrices and the…

最优化与控制 · 数学 2016-08-19 Yuan-Hua Ni , Ji-Feng Zhang , Miroslav Krstic

A computational method for the synthesis of time-optimal feedback control laws for linear nilpotent systems is proposed. The method is based on the use of the bang-bang theorem, which leads to a characterization of the time-optimal…

最优化与控制 · 数学 2026-05-20 Sara Bicego , Samuel Gue , Dante Kalise , Nelly Villamizar

In this article, two methods for solving mean-field type optimal control problems are proposed and investigated. The two methods are iterative methods: at each iteration, a Hamilton-Jacobi-Bellman equation is solved, for a terminal…

最优化与控制 · 数学 2017-03-30 Laurent Pfeiffer