中文
相关论文

相关论文: Sparse Stochastic Finite-State Controllers for POM…

200 篇论文

We propose a new approach to the problem of searching a space of policies for a Markov decision process (MDP) or a partially observable Markov decision process (POMDP), given a model. Our approach is based on the following observation: Any…

人工智能 · 计算机科学 2013-01-18 Andrew Y. Ng , Michael I. Jordan

The paper aims at the development of an apparatus for analysis and construction of near optimal solutions of singularly perturbed (SP) optimal controls problems (that is, problems of optimal control of SP systems) considered on the infinite…

最优化与控制 · 数学 2014-08-20 Vladimir Gaitsgory , Sergei Rossomakhine

Layered control is essential for managing complexity in large-scale systems, employing progressively coarser models at higher layers. While significant advances have been made for fully observable systems, the theoretical foundations of…

系统与控制 · 电气工程与系统科学 2026-04-15 Charis Stamouli , Anastasios Tsiamis , George J. Pappas

Recently discovered polyhedral structures of the value function for finite state-action discounted Markov decision processes (MDP) shed light on understanding the success of reinforcement learning. We investigate the value function polytope…

机器学习 · 计算机科学 2022-06-27 Yue Wu , Jesús A. De Loera

This paper investigates the limit behavior of Markov Decision Processes (MDPs) made of independent particles evolving in a common environment, when the number of particles goes to infinity. In the finite horizon case or with a discounted…

概率论 · 数学 2009-06-10 Nicolas Gast , Bruno Gaujal

This article proposes a new approach based on finite-horizon parameterizing manifolds (PMs) for the design of low-dimensional suboptimal controllers to optimal control problems of nonlinear partial differential equations (PDEs) of parabolic…

最优化与控制 · 数学 2014-11-19 Mickaël D. Chekroun , Honghu Liu

In ergodic singular stochastic control problems, a decision-maker can instantaneously adjust the evolution of a state variable using a control of bounded variation, with the goal of minimizing a long-term average cost functional. The cost…

最优化与控制 · 数学 2025-10-14 Alessandro Calvia , Federico Cannerozzi , Giorgio Ferrari

Many large MDPs can be represented compactly using a dynamic Bayesian network. Although the structure of the value function does not retain the structure of the process, recent work has shown that value functions in factored MDPs can often…

人工智能 · 计算机科学 2013-01-18 Daphne Koller , Ron Parr

We consider infinite-horizon stationary $\gamma$-discounted Markov Decision Processes, for which it is known that there exists a stationary optimal policy. Using Value and Policy Iteration with some error $\epsilon$ at each iteration, it is…

机器学习 · 计算机科学 2012-11-30 Bruno Scherrer , Boris Lesner

This paper focuses on learning a Constrained Markov Decision Process (CMDP) via general parameterized policies. We propose a Primal-Dual based Regularized Accelerated Natural Policy Gradient (PDR-ANPG) algorithm that uses entropy and…

机器学习 · 计算机科学 2026-05-04 Washim Uddin Mondal , Vaneet Aggarwal

Constrained decision-making is essential for designing safe policies in real-world control systems, yet simulated environments often fail to capture real-world adversities. We consider the problem of learning a policy that will maximize the…

机器学习 · 计算机科学 2026-02-10 Sourav Ganguly , Kishan Panaganti , Arnob Ghosh , Adam Wierman

Capturing uncertainty in models of complex dynamical systems is crucial to designing safe controllers. Stochastic noise causes aleatoric uncertainty, whereas imprecise knowledge of model parameters leads to epistemic uncertainty. Several…

系统与控制 · 电气工程与系统科学 2022-12-08 Thom Badings , Licio Romao , Alessandro Abate , Nils Jansen

This paper presents a constrained adaptive dynamic programming (CADP) algorithm to solve general nonlinear nonaffine optimal control problems with known dynamics. Unlike previous ADP algorithms, it can directly deal with problems with state…

系统与控制 · 电气工程与系统科学 2022-04-11 Jingliang Duan , Zhengyu Liu , Shengbo Eben Li , Qi Sun , Zhenzhong Jia , Bo Cheng

We establish the dual notions of scaling and saturation from geometric control theory in an infinite-dimensional setting. This generalization is applied to the low-mode control problem in a number of concrete nonlinear partial differential…

概率论 · 数学 2018-09-21 Nathan E. Glatt-Holtz , David P. Herzog , Jonathan C. Mattingly

We consider off-policy policy evaluation when the trajectory data are generated by multiple behavior policies. Recent work has shown the key role played by the state or state-action stationary distribution corrections in the infinite…

机器学习 · 计算机科学 2019-10-14 Xinyun Chen , Lu Wang , Yizhe Hang , Heng Ge , Hongyuan Zha

This paper looks at solving collaborative planning problems formalized as Decentralized POMDPs (Dec-POMDPs) by searching for Nash equilibria, i.e., situations where each agent's policy is a best response to the other agents' (fixed)…

人工智能 · 计算机科学 2021-09-21 Yang You , Vincent Thomas , Francis Colas , Olivier Buffet

Many control problems in environments that can be modeled as Markov decision processes (MDPs) concern infinite-time horizon specifications. The classical aim in this context is to compute a control policy that maximizes the probability of…

系统与控制 · 计算机科学 2017-05-03 Ruediger Ehlers , Salar Moarref , Ufuk Topcu

This paper studies a structured compound stochastic program (SP) involving multiple expectations coupled by nonconvex and nonsmooth functions. We present a successive convex-programming based sampling algorithm and establish its…

最优化与控制 · 数学 2021-05-25 Junyi Liu , Ying Cui , Jong-Shi Pang

This manuscript contains technical results related to a particular approach for the design of Model Predictive Control (MPC) laws. The approach, named "generalized" terminal state constraint, induces the recursive feasibility of the…

系统与控制 · 计算机科学 2013-07-16 Lorenzo Fagiano , Andrew R. Teel

We present Coordinated Proximal Policy Optimization (CoPPO), an algorithm that extends the original Proximal Policy Optimization (PPO) to the multi-agent setting. The key idea lies in the coordinated adaptation of step size during the…

人工智能 · 计算机科学 2021-11-09 Zifan Wu , Chao Yu , Deheng Ye , Junge Zhang , Haiyin Piao , Hankz Hankui Zhuo
‹ 上一页 1 8 9 10 下一页 ›