中文
相关论文

相关论文: Policy iteration for discrete-time systems with di…

200 篇论文

The paper deals with the stability analysis of time-delay reset control systems, for which the resetting law is assumed to satisfy a time-dependent condition. A stability analysis of the closed-loop system is performed based on an…

系统与控制 · 计算机科学 2016-03-09 M. A. Davó , F. Gouaisbaut , A. Baños , S. Tarbouriech , A. Seuret

In this paper, we consider the problem of optimizing the worst-case behavior of a partially observed system. All uncontrolled disturbances are modeled as finite-valued uncertain variables. Using the theory of cost distributions, we present…

最优化与控制 · 数学 2023-02-21 Aditya Dave , Nishanth Venkatesh , Andreas A. Malikopoulos

Distributionally robust policy learning aims to find a policy that performs well under the worst-case distributional shift, and yet most existing methods for robust policy learning consider the worst-case joint distribution of the covariate…

机器学习 · 计算机科学 2025-06-03 Jingyuan Wang , Zhimei Ren , Ruohan Zhan , Zhengyuan Zhou

Under non-exponential discounting, we develop a dynamic theory for stopping problems in continuous time. Our framework covers discount functions that induce decreasing impatience. Due to the inherent time inconsistency, we look for…

最优化与控制 · 数学 2017-03-13 Yu-Jui Huang , Adrien Nguyen-Huu

Motivated from Bertsekas' recent study on policy iteration (PI) for solving the problems of infinite-horizon discounted Markov decision processes (MDPs) in an on-line setting, we develop an off-line PI integrated with a multi-policy…

最优化与控制 · 数学 2021-12-07 Hyeong Soo Chang

Generating accurate runtime safety estimates for autonomous systems is vital to ensuring their continued proliferation. However, exhaustive reasoning about future behaviors is generally too complex to do at runtime. To provide scalable and…

计算机科学中的逻辑 · 计算机科学 2023-03-30 Matthew Cleaveland , Oleg Sokolsky , Insup Lee , Ivan Ruchkin

In this paper, we propose a new policy iteration algorithm to compute the value function and the optimal controls of continuous time stochastic control problems. The algorithm relies on successive approximations using linear-quadratic…

最优化与控制 · 数学 2024-09-09 Dylan Possamaï , Ludovic Tangpi

Learned action policies are increasingly popular in sequential decision-making, but suffer from a lack of safety guarantees. Recent work introduced a pipeline for testing the safety of such policies under initial-state and action-outcome…

人工智能 · 计算机科学 2026-03-17 Johannes Schmalz , Chaahat Jain

Many large MDPs can be represented compactly using a dynamic Bayesian network. Although the structure of the value function does not retain the structure of the process, recent work has shown that value functions in factored MDPs can often…

人工智能 · 计算机科学 2013-01-18 Daphne Koller , Ron Parr

In distributed model predictive control (DMPC), where a centralized optimization problem is solved in distributed fashion using dual decomposition, it is important to keep the number of iterations in the solution algorithm, i.e. the amount…

最优化与控制 · 数学 2013-07-11 Pontus Giselsson , Anders Rantzer

This paper presents a novel distributed model predictive control (MPC) formulation without terminal cost and a corresponding distributed synthesis approach for distributed linear discrete-time systems with coupled constraints. The proposed…

系统与控制 · 电气工程与系统科学 2026-05-28 Xiaoyu Liu , Dimos V. Dimarogonas , Changxin Liu , Azita Dabiri , Bart De Schutter

We construct control policies that ensure bounded variance of a noisy marginally stable linear system in closed-loop. It is assumed that the noise sequence is a mutually independent sequence of random vectors, enters the dynamics affinely,…

We consider challenging dynamic programming models where the associated Bellman equation, and the value and policy iteration algorithms commonly exhibit complex and even pathological behavior. Our analysis is based on the new notion of…

最优化与控制 · 数学 2016-09-13 Dimitri P. Bertsekas

We study computationally and statistically efficient reinforcement learning under the linear $Q^{\pi}$ realizability assumption, where any policy's $Q$-function is linear in a given state-action feature representation. Prior methods in this…

机器学习 · 计算机科学 2026-03-03 Yijing Ke , Zihan Zhang , Ruosong Wang

Recently, a novel class of Approximate Policy Iteration (API) algorithms have demonstrated impressive practical performance (e.g., ExIt from [2], AlphaGo-Zero from [27]). This new family of algorithms maintains, and alternately optimizes,…

机器学习 · 计算机科学 2019-04-09 Wen Sun , Geoffrey J. Gordon , Byron Boots , J. Andrew Bagnell

This paper introduces a new framework for analyzing the stability of discrete-time model predictive controllers acting on continuous-time systems. The proposed framework introduces the distinction between discretization time (used to…

系统与控制 · 电气工程与系统科学 2023-10-05 Yaashia Gautam , Marco M. Nicotra

Adaptive optimal control using value iteration initiated from a stabilizing control policy is theoretically analyzed in terms of stability of the system during the learning stage without ignoring the effects of approximation errors. This…

最优化与控制 · 数学 2017-10-25 Ali Heydari

We consider the problem of optimizing the steady state of a dynamical system in closed loop. Conventionally, the design of feedback optimization control laws assumes that the system is stationary. However, in reality, the dynamics of the…

最优化与控制 · 数学 2020-05-11 Sandeep Menta , Adrian Hauswirth , Saverio Bolognani , Gabriela Hug , Florian Dörfler

An optimal finite-time process drives a given initial distribution to a given final one in a given time at the lowest cost as quantified by total entropy production. We prove that for system with discrete states this optimal process…

统计力学 · 物理学 2023-12-15 Benedikt Remlein , Udo Seifert

We study the policy iteration algorithm (PIA) for entropy-regularized stochastic control problems on an infinite time horizon with a large discount rate, focusing on two main scenarios. First, we analyze PIA with bounded coefficients where…

最优化与控制 · 数学 2025-05-28 Hung Vinh Tran , Zhenhua Wang , Yuming Paul Zhang