English
Related papers

Related papers: A Modified Adaptive Data-Enabled Policy Optimizati…

200 papers

Modern alignment pipelines are increasingly replacing expensive human preference labels with evaluations from large language models (LLM-as-Judge). However, AI labels can be systematically biased compared to high-quality human feedback…

Machine Learning · Statistics 2026-02-10 Xintao Xia , Zhiqiu Xia , Linjun Zhang , Zhanrui Cai

Policy Optimization (PO) is one of the most popular methods in Reinforcement Learning (RL). Thus, theoretical guarantees for PO algorithms have become especially important to the RL community. In this paper, we study PO in adversarial MDPs…

Machine Learning · Computer Science 2023-05-16 Tal Lancewicki , Aviv Rosenberg , Dmitry Sotnikov

The design of optimal disturbance accommodation and servomechanism controllers with limited plant model information is considered in this paper. Their closed-loop performance are compared using a performance metric called competitive ratio…

Optimization and Control · Mathematics 2013-07-31 F. Farokhi , C. Langbort , K. H. Johansson

This paper studies the linear quadratic regulation (LQR) problem of unknown discrete-time systems via dynamic output feedback learning control. In contrast to the state feedback, the optimality of the dynamic output feedback control for…

Systems and Control · Electrical Eng. & Systems 2025-05-29 Kedi Xie , Martin Guay , Shimin Wang , Fang Deng , Maobin Lu

This paper investigates the problem of data-driven stabilization for linear discrete-time switched systems with unknown switching dynamics. In the absence of noise, a data-based state feedback stabilizing controller can be obtained by…

Systems and Control · Electrical Eng. & Systems 2023-11-21 Wenjie Liu , Yifei Li , Jian Sun , Gang Wang , Jie Chen

Data-driven predictive control promises model-free wave-dampening strategies for Connected and Autonomous Vehicles (CAVs) in mixed traffic flow. However, its performance relies on data quality, which suffers from unknown noise and…

Systems and Control · Electrical Eng. & Systems 2024-10-03 Shuai Li , Chaoyi Chen , Haotian Zheng , Jiawei Wang , Qing Xu , Keqiang Li

This paper implements the Deep Deterministic Policy Gradient (DDPG) algorithm for computing optimal policies for partially observable single-product periodic review inventory control problems with setup costs and backorders. The decision…

Optimization and Control · Mathematics 2025-07-29 Eugene Feinberg , Jefferson Huang , Pavlo Kasyanov , Thomas O'Neill

This paper presents a new control, namely additive-state-decomposition dynamic inversion stabilized control, that is used to stabilize a class of multi-input multi-output (MIMO) systems subject to nonparametric time-varying uncertainties…

Systems and Control · Computer Science 2020-03-10 Quan Quan , Guangxun Du , Kai-Yuan Cai

Policy gradient (PG) methods are the backbone of many reinforcement learning algorithms due to their good performance in policy optimization problems. As a gradient-based approach, PG methods typically rely on knowledge of the system…

Systems and Control · Electrical Eng. & Systems 2026-04-02 Bowen Song , Andrea Iannelli

Safe policy improvement (SPI) offers theoretical control over policy updates, yet existing guarantees largely concern offline, tabular reinforcement learning (RL). We study SPI in general online settings, when combined with world model and…

Machine Learning · Computer Science 2026-01-29 Florent Delgrange , Raphael Avalos , Willem Röpke

In the search for highly efficient decoders for short LDPC codes approaching maximum likelihood performance, a relayed decoding strategy, specifically activating the ordered statistics decoding process upon failure of a neural min-sum…

Information Theory · Computer Science 2024-03-26 Guangwen Li , Xiao Yu

DeepSeek-R1 has successfully enhanced Large Language Model (LLM) reasoning capabilities through its rule-based reward system. While it's a ''perfect'' reward system that effectively mitigates reward hacking, such reward functions are often…

Machine Learning · Computer Science 2025-10-27 Chenxing Wei , Jiarui Yu , Ying Tiffany He , Hande Dong , Yao Shu , Fei Yu

Test-time policy optimization enables large language models (LLMs) to adapt to distribution shifts by leveraging feedback from self-generated rollouts. However, existing methods rely on fixed-budget majority voting to estimate rewards,…

Machine Learning · Computer Science 2025-12-03 Youkang Wang , Jian Wang , Rubing Chen , Tianyi Zeng , Xiao-Yong Wei , Qing Li

When operating at their full capacity, quadrupedal robots can produce loud footstep noise, which can be disruptive in human-centered environments like homes, offices, and hospitals. As a result, balancing locomotion performance with noise…

Robotics · Computer Science 2025-03-10 Yuyou Zhang , Yihang Yao , Shiqi Liu , Yaru Niu , Changyi Lin , Yuxiang Yang , Wenhao Yu , Tingnan Zhang , Jie Tan , Ding Zhao

We consider the problem of learning control policies that optimize a reward function while satisfying constraints due to considerations of safety, fairness, or other costs. We propose a new algorithm, Projection-Based Constrained Policy…

Machine Learning · Computer Science 2020-10-08 Tsung-Yen Yang , Justinian Rosca , Karthik Narasimhan , Peter J. Ramadge

While direct policy optimization methods exist, pioneering LLMs are fine-tuned with reinforcement learning from human feedback (RLHF) to generate better responses under the supervision of a reward model learned from preference data. One…

Machine Learning · Computer Science 2025-06-10 Chuheng Zhang , Wei Shen , Li Zhao , Xuyun Zhang , Xiaolong Xu , Wanchun Dou , Jiang Bian

Hybrid Group Relative Policy Optimization (Hybrid GRPO) is a reinforcement learning framework that extends Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO) by incorporating empirical multi-sample action…

Machine Learning · Computer Science 2025-02-05 Soham Sane

For linear time-invariant systems, input-state data collected during an open-loop experiment can remedy the lack of knowledge of system parameters. However, such data do not contain information about other system uncertainties such as…

Optimization and Control · Mathematics 2025-10-02 Yongzhang Li , Amir Shakouri , M. Kanat Camlibel

Identifying unknown differential equations from a given set of discrete time dependent data is a challenging problem. A small amount of noise can make the recovery unstable, and nonlinearity and differential equations with varying…

Numerical Analysis · Mathematics 2019-04-09 Sung Ha Kang , Wenjing Liao , Yingjie Liu

Aligning large language models with pointwise absolute rewards has so far required online, on-policy algorithms such as PPO and GRPO. In contrast, simpler methods that can leverage offline or off-policy data, such as DPO and REBEL, are…

Machine Learning · Computer Science 2025-12-02 Simon Matrenok , Skander Moalla , Caglar Gulcehre