中文
相关论文

相关论文: An Adaptive Data-Enabled Policy Optimization Appro…

200 篇论文

This paper proposes modifications to the data-enabled policy optimization (DeePO) algorithm to mitigate state perturbations. DeePO is an adaptive, data-driven approach designed to iteratively compute a feedback gain equivalent to the…

系统与控制 · 电气工程与系统科学 2025-07-29 Mojtaba Kaheni , Niklas Persson , Vittorio De Iuliis , Costanzo Manes , Alessandro V. Papadopoulos

Data-enabled policy optimization (DeePO) is a newly proposed method to attack the open problem of direct adaptive LQR. In this work, we extend the DeePO framework to the linear quadratic tracking (LQT) with offline data. By introducing a…

系统与控制 · 电气工程与系统科学 2024-10-10 Shubo Kang , Feiran Zhao , Keyou You

Policy optimization (PO), an essential approach of reinforcement learning for a broad range of system classes, requires significantly more system data than indirect (identification-followed-by-control) methods or behavioral-based direct…

最优化与控制 · 数学 2023-09-18 Feiran Zhao , Florian Dörfler , Keyou You

Power electronic converters are becoming the main components of modern power systems due to the increasing integration of renewable energy sources. However, power converters may become unstable when interacting with the complex and…

系统与控制 · 电气工程与系统科学 2025-04-09 Feiran Zhao , Ruohan Leng , Linbin Huang , Huanhai Xin , Keyou You , Florian Dörfler

Direct data-driven design methods for the linear quadratic regulator (LQR) mainly use offline or episodic data batches, and their online adaptation has been acknowledged as an open problem. In this paper, we propose a direct adaptive method…

最优化与控制 · 数学 2024-10-07 Feiran Zhao , Florian Dörfler , Alessandro Chiuso , Keyou You

Model-free or learning-based control, in particular, reinforcement learning (RL), is expected to be applied for complex robotic tasks. Traditional RL requires a policy to be optimized is state-dependent, that means, the policy is a kind of…

机器学习 · 计算机科学 2022-08-09 Taisuke Kobayashi , Kenta Yoshizawa

Direct data-driven optimal control provides an elegant end-to-end paradigm, yet its real-time applicability is often hindered by the growing dimensionality of online decision variables. Recent breakthroughs, notably Data-EnablEd Policy…

系统与控制 · 电气工程与系统科学 2026-05-18 Shubo Kang , Keyou You

Federated Reinforcement Learning (FRL) has been deemed as a promising solution for intelligent decision-making in the era of Artificial Internet of Things. However, existing FRL approaches often entail repeated interactions with the…

机器学习 · 计算机科学 2024-05-30 Sheng Yue , Zerui Qin , Xingyuan Hua , Yongheng Deng , Ju Ren

Federated learning (FL) has emerged as a solution to deal with the risk of privacy leaks in machine learning training. This approach allows a variety of mobile devices to collaboratively train a machine learning model without sharing the…

机器学习 · 计算机科学 2022-12-01 Young Geun Kim , Carole-Jean Wu

We are motivated by the real challenges presented in a human-robot system to develop new designs that are efficient at data level and with performance guarantees such as stability and optimality at systems level. Existing…

系统与控制 · 电气工程与系统科学 2021-01-19 Xiang Gao , Jennie Si , Yue Wen , Minhan Li , He , Huang

This study proposes a delay-compensated feedback controller based on proximal policy optimization (PPO) reinforcement learning to stabilize traffic flow in the congested regime by manipulating the time-gap of adaptive cruise…

人工智能 · 计算机科学 2023-01-18 Shurong Mo , Nailong Wu , Jie Qi , Anqi Pan , Zhiguang Feng , Huaicheng Yan , Yueying Wang

Data-enabled Predictive Control (DeePC) has recently gained the spotlight as an easy-to-use control technique that allows for constraint handling while relying on raw data only. Initially proposed for linear time-invariant systems, several…

系统与控制 · 电气工程与系统科学 2025-04-14 Gianluca Giacomelli , Simone Formentin , Victor G. Lopez , Matthias A. Müller , Valentina Breschi

This note proposes a data-driven output-feedback stabilizing policy iteration for unknown linear discrete-time systems with unmeasurable states. Existing policy iteration methods for optimal control must start from a stabilizing control…

系统与控制 · 电气工程与系统科学 2025-12-01 Dongdong Li , Jiuxiang Dong

In this paper, we study a data-enabled predictive control (DeePC) algorithm applied to unknown stochastic linear time-invariant systems. The algorithm uses noise-corrupted input/output data to predict future trajectories and compute optimal…

最优化与控制 · 数学 2019-11-04 Jeremy Coulson , John Lygeros , Florian Dörfler

Offline reinforcement learning (RL) aims to learn decision policies from a fixed batch of logged transitions, without additional environment interaction. Despite remarkable empirical progress, offline RL remains fragile under distribution…

统计方法学 · 统计学 2026-03-16 Debashis Chatterjee

Autonomous bicycles offer a promising agile solution for urban mobility and last-mile logistics. However, conventional control strategies often struggle with underactuated nonlinear dynamics, suffering from sensitivity to model mismatches…

机器人学 · 计算机科学 2026-05-05 Gelu Liu , Teng Wang , Zhijie Wu , Junliang Wu , Songyuan Li , Xiangwei Zhu

Synthetic data is central to data-efficient Dyna-style model-based reinforcement learning, but it can also degrade performance. We study this failure in Model-Based Policy Optimization (MBPO), which performs actor-critic updates using…

机器学习 · 计算机科学 2026-05-08 Brett Barkley , David Fridovich-Keil

Data-enabled predictive control (DeePC) has garnered significant attention for its ability to achieve safe, data-driven optimal control without relying on explicit system models. Traditional DeePC methods use pre-collected input/output…

系统与控制 · 电气工程与系统科学 2024-07-24 Amin Vahidi-Moghaddam , Kaixiang Zhang , Xunyuan Yin , Vaibhav Srivastava , Zhaojian Li

This paper considers a distributed reinforcement learning problem for decentralized linear quadratic control with partial state observations and local costs. We propose a Zero-Order Distributed Policy Optimization algorithm (ZODPO) that…

系统与控制 · 电气工程与系统科学 2020-10-26 Yingying Li , Yujie Tang , Runyu Zhang , Na Li

Post-training paradigms for Large Language Models (LLMs), primarily Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), face a fundamental dilemma: SFT provides stability (low variance) but suffers from high fitting bias, while RL…

机器学习 · 计算机科学 2026-04-13 Taojie Zhu , Dongyang Xu , Ding Zou , Sen Zhao , Qiaobo Hao , Zhiguo Yang , Yonghong He
‹ 上一页 1 2 3 10 下一页 ›