中文
相关论文

相关论文: An Adaptive Data-Enabled Policy Optimization Appro…

200 篇论文

Modern alignment pipelines are increasingly replacing expensive human preference labels with evaluations from large language models (LLM-as-Judge). However, AI labels can be systematically biased compared to high-quality human feedback…

机器学习 · 统计学 2026-02-10 Xintao Xia , Zhiqiu Xia , Linjun Zhang , Zhanrui Cai

Model predictive control (MPC) is increasingly being considered for control of fast systems and embedded applications. However, the MPC has some significant challenges for such systems. Its high computational complexity results in high…

系统与控制 · 电气工程与系统科学 2024-10-28 Eivind Bøhn , Sebastien Gros , Signe Moe , Tor Arne Johansen

A modification to the ${\cal L}_1$ control framework for uncertain systems with actuator delay is presented. Specifically, a time delay is introduced in the control input of the state predictor to compensate for the destabilizing effect of…

最优化与控制 · 数学 2022-09-27 Kim-Doang Nguyen , Harry Dankowicz

A reliable controller is critical and essential for the execution of safe and smooth maneuvers of an autonomous vehicle.The controller must be robust to external disturbances, such as road surface, weather, and wind conditions, and so on.It…

机器人学 · 计算机科学 2019-05-01 Tianyu Shi , Pin Wang , Ching-Yao Chan , Chonghao Zou

Artificial time delay controller was conceptualised for nonlinear systems to reduce dependency on precise system modelling unlike the conventional adaptive and robust control strategies. In this approach unknown dynamics is compensated by…

机器人学 · 计算机科学 2024-09-04 Swati Dantu

Reinforcement learning from human feedback (RLHF) typically assumes a static or non-strategic reward model (RM). In iterative deployment, however, the policy generates the data on which the RM is retrained, creating a feedback loop.…

机器学习 · 计算机科学 2026-05-07 Etienne Gauthier , Francis Bach , Michael I. Jordan

Direct Preference Optimization (DPO) has emerged as a predominant alignment method for diffusion models, facilitating off-policy training without explicit reward modeling. However, its reliance on large-scale, high-quality human preference…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Khiem Pham , Quang Nguyen , Tung Nguyen , Jingsen Zhu , Michele Santacatterina , Dimitris Metaxas , Ramin Zabih

Deep Reinforcement Learning (DRL) has experienced significant advancements in recent years and has been widely used in many fields. In DRL-based robotic policy learning, however, current de facto policy parameterization is still…

机器人学 · 计算机科学 2026-03-13 Diyuan Shi , Yiqi Tang , Zifeng Zhuang , Donglin Wang

In optimal control problem, policy iteration (PI) is a powerful reinforcement learning (RL) tool used for designing optimal controller for the linear systems. However, the need for an initial stabilizing control policy significantly limits…

最优化与控制 · 数学 2024-11-13 Zhen Pang , Shengda Tang , Jun Cheng , Shuping He

We propose the application of Koopman operator theory for the design of stabilizing feedback controller for a nonlinear control system. The proposed approach is data-driven and relies on the use of time-series data generated from the…

最优化与控制 · 数学 2019-01-24 Bowen Huang , Xu Ma , Umesh Vaidya

This paper introduces an analytical framework for the derivation of hybrid equations of motion of a flexible quadrotor. This approach helps obtain rigid and elastic equations of motion simultaneously, in a decoupled form, which facilitates…

系统与控制 · 电气工程与系统科学 2020-12-11 Emre Eraslan , Yildiray Yildiz

Motion planning and control are two core components of the robotic systems autonomy stack. The standard approach to combine these methodologies comprises an offline/open-loop stage, planning, that designs a feasible and safe trajectory to…

系统与控制 · 电气工程与系统科学 2023-10-23 Tianqi Zheng , John W. Simpson-Porco , Enrique Mallada

This paper develops an adaptive observation-based efficient reinforcement learning (RL) approach for systems with uncertain drift dynamics. A novel concurrent learning adaptive extended observer (CL-AEO) is first designed to jointly…

动力系统 · 数学 2020-11-25 Maopeng Ran , Lihua Xie

Federated learning (FL) is a machine learning model that preserves data privacy in the training process. Specifically, FL brings the model directly to the user equipments (UEs) for local training, where an edge server periodically collects…

信息论 · 计算机科学 2019-11-01 Howard H. Yang , Ahmed Arafa , Tony Q. S. Quek , H. Vincent Poor

Model Predictive Control (MPC)-based Reinforcement Learning (RL) offers a structured and interpretable alternative to Deep Neural Network (DNN)-based RL methods, with lower computational complexity and greater transparency. However,…

系统与控制 · 电气工程与系统科学 2025-07-15 Hossein Nejatbakhsh Esfahani , Javad Mohammadpour Velni

Direct Preference Optimization (DPO) and its variants have become standard for aligning Large Language Models due to their simplicity and offline stability. However, we identify two fundamental limitations. First, the optimal policy depends…

人工智能 · 计算机科学 2026-02-10 Yu Li , Tian Lan , Zhengling Qi

Federated learning (FL) is a promising distributed paradigm, eliminating the need for data sharing but facing challenges from data heterogeneity. Personalized parameter generation through a hypernetwork proves effective, yet existing…

计算机视觉与模式识别 · 计算机科学 2023-11-22 Ziyuan Yang , Zerui Shao , Huijie Huangfu , Hui Yu , Andrew Beng Jin Teoh , Xiaoxiao Li , Hongming Shan , Yi Zhang

We present a trajectory planning and control architecture for bipedal locomotion at a variety of speeds on a highly underactuated and compliant bipedal robot. A library of compliant walking trajectories are planned offline, and stored as…

机器人学 · 计算机科学 2020-10-20 Jenna Reher , Aaron D. Ames

Model-free and reinforcement learning-based adaptive filtering methods are gaining traction for denoising in dynamic, non-stationary environments such as wireless signal channels. Traditional filters like LMS, RLS, Wiener, and Kalman are…

信号处理 · 电气工程与系统科学 2025-06-10 Abdullah Burkan Bereketoglu

Motion planning for autonomous driving (AD) faces a critical trade-off. While traditional rule-based pipelines offer verifiable safety and interpretability, they often fail to generalize in complex scenarios. Conversely, emerging…

机器人学 · 计算机科学 2026-05-29 Jia Hu , Yang Chang , Haoran Wang