中文
相关论文

相关论文: Convergence of machine learning methods for feedba…

200 篇论文

An intelligent Real-Time Sensing (RTS) system must continuously acquire, update, integrate, and apply knowledge to adapt to real-world dynamics. Managing distributed intelligence in this context requires Federated Continual Learning (FCL).…

Automated machine learning (AutoML) systems aim to enable training machine learning (ML) models for non-ML experts. A shortcoming of these systems is that when they fail to produce a model with high accuracy, the user has no path to improve…

机器学习 · 计算机科学 2021-02-23 Behnaz Arzani , Kevin Hsieh , Haoxian Chen

We establish the convergence of the unified two-timescale Reinforcement Learning (RL) algorithm presented in a previous work by Angiuli et al. This algorithm provides solutions to Mean Field Game (MFG) or Mean Field Control (MFC) problems…

最优化与控制 · 数学 2024-05-02 Andrea Angiuli , Jean-Pierre Fouque , Mathieu Laurière , Mengrui Zhang

Alignment with human preferences is commonly framed using a universal reward function, even though human preferences are inherently heterogeneous. We formalize this heterogeneity by introducing user types and examine the limits of the…

人工智能 · 计算机科学 2025-02-25 Ali Shirali , Arash Nasr-Esfahany , Abdullah Alomar , Parsa Mirtaheri , Rediet Abebe , Ariel Procaccia

In this paper, we propose a generic framework for devising an adaptive approximation scheme for value function approximation in reinforcement learning, which introduces multiscale approximation. The two basic ingredients are multiresolution…

机器学习 · 计算机科学 2019-08-26 Tao Li , Quanyan Zhu

We stabilize the flow past a cluster of three rotating cylinders, the fluidic pinball, with automated gradient-enriched machine learning algorithms. The control laws command the rotation speed of each cylinder in an open- and closed-loop…

流体动力学 · 物理学 2022-06-16 Guy Y. Cornejo Maceda , Yiqing Li , François Lusseyran , Marek Morzyński , Bernd R. Noack

Federated learning is highly valued due to its high-performance computing in distributed environments while safeguarding data privacy. To address resource heterogeneity, researchers have proposed a semi-asynchronous federated learning…

分布式、并行与集群计算 · 计算机科学 2024-05-28 Yunbo Li , Jiaping Gui , Yue Wu

We present DRESS, a large vision language model (LVLM) that innovatively exploits Natural Language feedback (NLF) from Large Language Models to enhance its alignment and interactions by addressing two key limitations in the state-of-the-art…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Yangyi Chen , Karan Sikka , Michael Cogswell , Heng Ji , Ajay Divakaran

We show two average-reward off-policy control algorithms, Differential Q-learning (Wan, Naik, & Sutton 2021a) and RVI Q-learning (Abounadi Bertsekas & Borkar 2001), converge in weakly communicating MDPs. Weakly communicating MDPs are the…

机器学习 · 计算机科学 2022-11-08 Yi Wan , Richard S. Sutton

We introduce Feasible Learning (FL), a sample-centric learning paradigm where models are trained by solving a feasibility problem that bounds the loss for each training sample. In contrast to the ubiquitous Empirical Risk Minimization (ERM)…

We propose an improved convergence analysis technique that characterizes the distributed learning paradigm of federated learning (FL) with imperfect/noisy uplink and downlink communications. Such imperfect communication scenarios arise in…

机器学习 · 计算机科学 2023-07-17 Antesh Upadhyay , Abolfazl Hashemi

Federated Learning (FL) is a collaborative machine learning (ML) framework that combines on-device training and server-based aggregation to train a common ML model among distributed agents. In this work, we propose an asynchronous FL design…

机器学习 · 计算机科学 2025-12-04 Chung-Hsuan Hu , Zheng Chen , Erik G. Larsson

Reinforcement Learning from Human Feedback (RLHF) is commonly employed to tailor models to human preferences, especially to improve the safety of outputs from large language models (LLMs). Traditionally, this method depends on selecting…

计算与语言 · 计算机科学 2025-01-29 Xiaomin Li , Mingye Gao , Zhiwei Zhang , Jingxuan Fan , Weiyu Li

Reinforcement Learning from Human Feedback (RLHF) is popular in large language models (LLMs), whereas traditional Reinforcement Learning (RL) often falls short. Current autonomous driving methods typically utilize either human feedback in…

人工智能 · 计算机科学 2024-10-10 Yuan Sun , Navid Salami Pargoo , Peter J. Jin , Jorge Ortiz

This survey provides an overview of combining Federated Learning (FL) and control to enhance adaptability, scalability, generalization, and privacy in (nonlinear) control applications. Traditional control methods rely on controller design…

机器学习 · 计算机科学 2024-11-15 Jakob Weber , Markus Gurtner , Amadeus Lobe , Adrian Trachte , Andreas Kugi

We consider a mean-field control problem in which admissible controls are required to be adapted to the common noise filtration. The main objective is to show how the mean-field control problem can be approximates by time consistent…

最优化与控制 · 数学 2025-09-19 Bruno Bouchard , Xiaolu Tan

This paper studies optimal control under the average-reward/cost criterion for deterministic linear systems. We derive the value function and optimal policy, and propose an approximate solution using Model Predictive Control to enable…

最优化与控制 · 数学 2025-07-08 Duc Cuong Nguyen

Software fault prediction (SFP) is a critical task in software engineering, enabling early identification of faults in modules to improve software quality and reduce maintenance costs. This research investigates the combined effects of…

Safety constraints and optimality are important, but sometimes conflicting criteria for controllers. Although these criteria are often solved separately with different tools to maintain formal guarantees, it is also common practice in…

系统与控制 · 电气工程与系统科学 2024-06-10 Pierre-François Massiani , Steve Heim , Friedrich Solowjow , Sebastian Trimpe

In this study, we delve into Federated Reinforcement Learning (FedRL) in the context of value-based agents operating across diverse Markov Decision Processes (MDPs). Existing FedRL methods typically aggregate agents' learning by averaging…

机器学习 · 计算机科学 2024-04-18 Hei Yi Mak , Flint Xiaofeng Fan , Luca A. Lanzendörfer , Cheston Tan , Wei Tsang Ooi , Roger Wattenhofer