中文
相关论文

相关论文: On Entropy Regularized Path Integral Control for T…

200 篇论文

Model predictive control (MPC) has established itself as the primary methodology for constrained control, enabling general-purpose robot autonomy in diverse real-world scenarios. However, for most problems of interest, MPC relies on the…

机器人学 · 计算机科学 2024-11-01 Davide Celestini , Daniele Gammelli , Tommaso Guffanti , Simone D'Amico , Elisa Capello , Marco Pavone

The entropy regularization is inspired by information entropy from machine learning and the ideas of exploration and exploitation in reinforcement learning, which appears in the control problem to design an approximating algorithm for the…

最优化与控制 · 数学 2024-11-21 Ziyue Chen , Qi Zhang

We present a data-driven optimal control framework that can be viewed as a generalization of the path integral (PI) control approach. We find iterative feedback control laws without parameterization based on probabilistic representation of…

系统与控制 · 计算机科学 2016-02-02 Yunpeng Pan , Evangelos A. Theodorou , Michail Kontitsis

Policy optimization is an effective reinforcement learning approach to solve continuous control tasks. Recent achievements have shown that alternating online and offline optimization is a successful choice for efficient trajectory reuse.…

机器学习 · 计算机科学 2018-11-01 Alberto Maria Metelli , Matteo Papini , Francesco Faccio , Marcello Restelli

In this paper, we explore a scenario where a sender provides an information policy and a receiver, upon observing a realization of this policy, decides whether to take a particular action, such as making a purchase. The sender's objective…

数值分析 · 数学 2024-12-13 Jorge Justiniano , Andreas Kleiner , Benny Moldovanu , Martin Rumpf , Philipp Strack

Discrete-time stochastic optimal control remains a challenging problem for general, nonlinear systems under significant uncertainty, with practical solvers typically relying on the certainty equivalence assumption, replanning and/or…

系统与控制 · 电气工程与系统科学 2021-03-12 Joe Watson , Jan Peters

This paper proposes an iterative distributionally robust model predictive control (MPC) scheme to solve a risk-constrained infinite-horizon optimal control problem. In each iteration, the algorithm generates a trajectory from the starting…

最优化与控制 · 数学 2023-08-23 Alireza Zolanvari , Ashish Cherukuri

We address the generic problem of optimal quantum state preparation for open quantum systems. It is well known that open quantum systems can be simulated by quantum trajectories described by a stochastic Schr\"odinger equation. In this…

量子物理 · 物理学 2025-01-31 Aarón Villanueva , Hilbert Kappen

Training LLM agents in multi-turn environments with sparse rewards, where completing a single task requires 30+ turns of interaction within an episode, presents a fundamental challenge for reinforcement learning. We identify a critical…

机器学习 · 计算机科学 2026-02-11 Wujiang Xu , Wentian Zhao , Zhenting Wang , Yu-Jhe Li , Can Jin , Mingyu Jin , Kai Mei , Kun Wan , Dimitris N. Metaxas

In this paper, the solvability of the Inverse Optimal Control (IOC) problem based on two existing minimum principal methods, is analysed. The aim of this work is to answer the question regarding what kinds of trajectories, that is depending…

最优化与控制 · 数学 2024-03-15 Afreen Islam , Guido Herrmann , Joaquin Carrasco

We consider a class of integer linear programs (IPs) that arise as discretizations of trust-region subproblems of a trust-region algorithm for the solution of control problems, where the control input is an integer-valued function on a…

最优化与控制 · 数学 2022-06-06 Marvin Severitt , Paul Manns

Optimal transport (OT) defines a powerful framework to compare probability distributions in a geometrically faithful way. However, the practical impact of OT is still limited because of its computational burden. We propose a new class of…

最优化与控制 · 数学 2016-05-30 Genevay Aude , Marco Cuturi , Gabriel Peyré , Francis Bach

We address the role of noise and the issue of efficient computation in stochastic optimal control problems. We consider a class of non-linear control problems that can be formulated as a path integral and where the noise plays the role of…

计算物理 · 物理学 2009-11-10 H. J. Kappen

We study the trajectory optimization problem under chance constraints for continuous-time stochastic systems. To address chance constraints imposed on the entire stochastic trajectory, we propose a framework based on the set erosion…

最优化与控制 · 数学 2025-04-08 Zishun Liu , Liqian Ma , Yongxin Chen

This paper explores continuous-time and state-space optimal stopping problems from a reinforcement learning perspective. We begin by formulating the stopping problem using randomized stopping times, where the decision maker's control is…

最优化与控制 · 数学 2026-03-12 Jodi Dianetti , Giorgio Ferrari , Renyuan Xu

This paper presents an algorithm to apply nonlinear control design approaches in the case of stochastic systems with partial state observation. Deterministic nonlinear control approaches are formulated under the assumption of full state…

系统与控制 · 电气工程与系统科学 2023-09-19 Mohammad S. Ramadan , Mohammad Alsuwaidan , Ahmed Atallah , Sylvia Herbert

We present a sampling-based Model Predictive Control (MPC) method that implements Model Predictive Path Integral (MPPI) as an \emph{Ising machine}, suitable for novel forms of probabilistic computing. By expressing the control problem as a…

系统与控制 · 电气工程与系统科学 2025-12-18 Lorin Werthen-Brabants , Pieter Simoens

Reinforcement Learning (RL) algorithms have shown tremendous success in simulation environments, but their application to real-world problems faces significant challenges, with safety being a major concern. In particular, enforcing…

机器学习 · 计算机科学 2024-06-19 Weiye Zhao , Rui Chen , Yifan Sun , Tianhao Wei , Changliu Liu

In this paper we investigate the convergence of the Policy Iteration Algorithm (PIA) for a class of general continuous-time entropy-regularized stochastic control problems. In particular, instead of employing sophisticated PDE estimates for…

最优化与控制 · 数学 2025-04-24 Jin Ma , Gaozhan Wang , Jianfeng Zhang

Language models (LMs) are trained on billions of tokens in an attempt to recover the true language distribution. Still, vanilla random sampling from LMs yields low quality generations. Decoding algorithms attempt to restrict the LM…

机器学习 · 计算机科学 2026-01-06 Kareem Ahmed , Sameer Singh