English
Related papers

Related papers: Convergence of Policy Iteration for Entropy-Regula…

200 papers

In this paper, we consider a class of continuous-time, continuous-space stochastic optimal control problems. Building upon recent advances in Markov chain approximation methods and sampling-based algorithms for deterministic path planning,…

Robotics · Computer Science 2012-02-27 Vu Anh Huynh , Sertac Karaman , Emilio Frazzoli

This paper aims to establish an entropy-regularized value-based reinforcement learning method that can ensure the monotonic improvement of policies at each policy update. Unlike previously proposed lower-bounds on policy improvement in…

Machine Learning · Computer Science 2020-08-26 Lingwei Zhu , Takamitsu Matsubara

Coordination of distributed agents is required for problems arising in many areas, including multi-robot systems, networking and e-commerce. As a formal framework for such problems, we use the decentralized partially observable Markov…

Artificial Intelligence · Computer Science 2014-01-16 Daniel S. Bernstein , Christopher Amato , Eric A. Hansen , Shlomo Zilberstein

Policy iteration (PI) is a recursive process of policy evaluation and improvement for solving an optimal decision-making/control problem, or in other words, a reinforcement learning (RL) problem. PI has also served as the fundamental for…

Artificial Intelligence · Computer Science 2021-04-06 Jaeyoung Lee , Richard S. Sutton

Environmental management optimizing a long-run objective is an ergodic control problem whose resolution can be achieved by solving an associated non-local Hamilton-Jacobi-Bellman (HJB) equation having an effective Hamiltonian. Focusing on…

Optimization and Control · Mathematics 2022-05-11 Hidekazu Yoshioka , Motoh Tsujimura , Yuta Yaegashi

The classical Dynamic Programming (DP) approach to optimal control problems is based on the characterization of the value function as the unique viscosity solution of a Hamilton-Jacobi-Bellman (HJB) equation. The DP scheme for the numerical…

Numerical Analysis · Mathematics 2019-04-15 Alessandro Alla , Maurizio Falcone , Luca Saluzzi

We study singular perturbations of a class of two-scale stochastic control systems with unbounded data. The assumptions are designed to cover some relaxation problems for deep neural networks. We construct effective Hamiltonian and initial…

Optimization and Control · Mathematics 2023-03-29 Martino Bardi , Hicham Kouhkouh

Bounded policy iteration is an approach to solving infinite-horizon POMDPs that represents policies as stochastic finite-state controllers and iteratively improves a controller by adjusting the parameters of each node using linear…

Artificial Intelligence · Computer Science 2012-06-18 Eric A. Hansen

This paper studies an optimal stochastic impulse control problem in a finite horizon with a decision lag, by which we mean that after an impulse is made, a fixed number units of time has to be elapsed before the next impulse is allowed to…

Optimization and Control · Mathematics 2021-02-09 Chang Li , Jiongmin Yong

In this paper, we prove both necessary and sufficient maximum principles for infinite horizon discounted control problems of stochastic Volterra integral equations with finite delay and a convex control domain. The corresponding adjoint…

Optimization and Control · Mathematics 2023-03-15 Yushi Hamaguchi

An optimal control problem is considered for a stochastic differential equation containing a state-dependent regime switching, with a recursive cost functional. Due to the non-exponential discounting in the cost functional, the problem is…

Optimization and Control · Mathematics 2017-12-29 Hongwei Mei , Jiongmin Yong

Regularization of control policies using entropy can be instrumental in adjusting predictability of real-world systems. Applications benefiting from such approaches range from, e.g., cybersecurity, which aims at maximal unpredictability, to…

Systems and Control · Electrical Eng. & Systems 2026-02-18 Menno van Zutphen , Giannis Delimpaltadakis , Maurice Heemels , Duarte Antunes

This paper proposes and analyzes two new policy learning methods: regularized policy gradient (RPG) and iterative policy optimization (IPO), for a class of discounted linear-quadratic control (LQC) problems over an infinite time horizon…

Optimization and Control · Mathematics 2025-10-08 Xin Guo , Xinyu Li , Renyuan Xu

Reinforcement learning based adaptive/approximate dynamic programming (ADP) is a powerful technique to determine an approximate optimal controller for a dynamical system. These methods bypass the need to analytically solve the nonlinear…

Optimization and Control · Mathematics 2018-05-24 Xuefeng Bao , Zhi-Hong Mao , Nitin Sharma

Following the recent resurgence in establishing linear control theoretic benchmarks for reinforcement leaning (RL)-based policy optimization (PO) for complex dynamical systems with continuous state and action spaces, an optimal control…

Systems and Control · Electrical Eng. & Systems 2023-06-30 Leilei Cui , Lekan Molu

We use the technique of information relaxation to develop a duality-driven iterative approach to obtaining and improving confidence interval estimates for the true value of finite-horizon stochastic dynamic programming problems. We show…

Optimization and Control · Mathematics 2020-07-29 Nan Chen , Xiang Ma , Yanchu Liu , Wei Yu

Solving the Hamilton-Jacobi-Bellman equation is important in many domains including control, robotics and economics. Especially for continuous control, solving this differential equation and its extension the Hamilton-Jacobi-Isaacs…

Robotics · Computer Science 2021-10-06 Michael Lutter , Boris Belousov , Shie Mannor , Dieter Fox , Animesh Garg , Jan Peters

We study the exploratory Hamilton--Jacobi--Bellman (HJB) equation arising from the entropy-regularized exploratory control problem, which was formulated by Wang, Zariphopoulou and Zhou (J. Mach. Learn. Res., 21, 2020) in the context of…

Optimization and Control · Mathematics 2021-09-22 Wenpin Tang , Paul Yuming Zhang , Xun Yu Zhou

In this paper we study an optimization problem in which the control is information, more precisely, the control is a $\sigma$-algebra or a filtration. In a dynamic setting, we establish the dynamic programming principle and the law…

Optimization and Control · Mathematics 2026-03-31 Zihao Gu , Jianfeng Zhang

We study high-dimensional stochastic optimal control problems in which many agents cooperate to minimize a convex cost functional. We consider both the full-information problem, in which each agent observes the states of all other agents,…

Probability · Mathematics 2023-01-10 Joe Jackson , Daniel Lacker
‹ Prev 1 3 4 5 6 7 10 Next ›