中文
相关论文

相关论文: A Switching System Theory of Q-Learning with Linea…

200 篇论文

Q-learning is a popular reinforcement learning algorithm. This algorithm has however been studied and analysed mainly in the infinite horizon setting. There are several important applications which can be modeled in the framework of finite…

机器学习 · 计算机科学 2022-08-09 Vivek VP , Dr. Shalabh Bhatnagar

Offline reinforcement learning leverages large datasets to train policies without interactions with the environment. The learned policies may then be deployed in real-world settings where interactions are costly or dangerous. Current…

机器学习 · 计算机科学 2022-06-29 Matthias Weissenbacher , Samarth Sinha , Animesh Garg , Yoshinobu Kawahara

We consider the problem of joint learning of multiple linear dynamical systems. This has received significant attention recently under different types of assumptions on the model parameters. The setting we consider involves a collection of…

最优化与控制 · 数学 2025-06-04 Hemant Tyagi

Q-learning is a regression-based approach that is widely used to formalize the development of an optimal dynamic treatment strategy. Finite dimensional working models are typically used to estimate certain nuisance parameters, and…

统计方法学 · 统计学 2020-03-30 Ashkan Ertefaie , James R. McKay , David Oslin , Robert L. Strawderman

We present an approach to identify a quasi Linear Parameter Varying (qLPV) model of a plant, with the qLPV model guaranteed to admit a robust control invariant (RCI) set. It builds upon the concurrent synthesis framework presented in [1],…

最优化与控制 · 数学 2025-05-13 Sampath Kumar Mulagaleti , Alberto Bemporad

We give an efficient algorithm for learning a binary function in a given class C of bounded VC dimension, with training data distributed according to P and test data according to Q, where P and Q may be arbitrary distributions over X. This…

机器学习 · 计算机科学 2021-02-17 Adam Kalai , Varun Kanade

As a primary contribution, we present a convergence theorem for stochastic iterations, and in particular, Q-learning iterates, under a general, possibly non-Markovian, stochastic environment. Our conditions for convergence involve an…

最优化与控制 · 数学 2024-03-05 Ali Devran Kara , Serdar Yuksel

The field of quickest change detection (QCD) concerns design and analysis of algorithms to estimate in real time the time at which an important event takes place, and identify properties of the post-change behavior. It is shown in this…

最优化与控制 · 数学 2024-09-16 Austin Cooper , Sean Meyn

Large-scale multi-agent systems are often deployed across wide geographic areas, where agents interact with heterogeneous environments. There is an emerging interest in understanding the role of heterogeneity in the performance of the…

机器学习 · 计算机科学 2026-05-18 Leo Muxing Wang , Pengkun Yang , Lili Su

Motivated by the growing use of artificial intelligence (AI) tools in control design, this paper analyses the intersection between results from gradient methods for the model-free linear quadratic regulator (LQR), and linear feedforward…

系统与控制 · 电气工程与系统科学 2025-05-27 Arthur Castello B. de Oliveira , Milad Siami , Eduardo D. Sontag

Pretrained on large-scale and diverse datasets, VLA models demonstrate strong generalization and adaptability as general-purpose robotic policies. However, Supervised Fine-Tuning (SFT), which serves as the primary mechanism for adapting…

机器人学 · 计算机科学 2026-05-19 Yuan Liu , Haoran Li , Shuai Tian , Yuxing Qin , Yuhui Chen , Yupeng Zheng , Yongzhen Huang , Dongbin Zhao

We propose a theoretical framework for an adaptive learning rate policy for the Mean Absolute Error loss function and Quantile loss function and evaluate its effectiveness for regression tasks. The framework is based on the theory of…

In this paper, we present a framework to understand the convergence of commonly used Q-learning reinforcement learning algorithms in practice. Two salient features of such algorithms are: (i)~the Q-table is recursively updated using an…

机器学习 · 计算机科学 2025-09-04 Amit Sinha , Matthieu Geist , Aditya Mahajan

The aim of this paper is to introduce a quantum fusion mechanism for multimodal learning and to establish its theoretical and empirical potential. The proposed method, called the Quantum Fusion Layer (QFL), replaces classical fusion schemes…

量子物理 · 物理学 2025-10-09 Tuyen Nguyen , Trong Nghia Hoang , Phi Le Nguyen , Hai L. Vu , Truong Cong Thang

SARSA, a classical on-policy control algorithm for reinforcement learning, is known to chatter when combined with linear function approximation: SARSA does not diverge but oscillates in a bounded region. However, little is known about how…

机器学习 · 计算机科学 2023-05-16 Shangtong Zhang , Remi Tachet , Romain Laroche

Due to technological advances in the field of radio technology and its availability, the number of interference signals in the radio spectrum is continuously increasing. Interference signals must be detected in a timely fashion, in order to…

系统与控制 · 电气工程与系统科学 2023-07-13 Tobias Braun , Tobias Korzyzkowske , Larissa Putzar , Jan Mietzner , Peter A. Hoeher

Density functional theory (DFT) and linear-response time-dependent density functional theory (LR-TDDFT) rely on an exchange-correlation (xc) approximation that provides not only energy but also its functional derivatives that enter the…

化学物理 · 物理学 2026-04-08 Xiaoyu Zhang

This paper addresses the stabilisation of discrete-time switching linear systems (DTSSs) with control inputs under arbitrary switching, based on the existence of a common quadratic Lyapunov function (CQLF). The authors have begun a line of…

系统与控制 · 计算机科学 2011-09-15 Hernan Haimovich , Julio H. Braslavsky

Shuffled linear regression (SLR) seeks to estimate latent features through a linear transformation, complicated by unknown permutations in the measurement dimensions. This problem extends traditional least-squares (LS) and Least Absolute…

统计理论 · 数学 2025-04-17 Hang Liu , Anna Scaglione

Detection of phase transitions is a critical task in statistical physics, traditionally pursued through analytic methods and direct numerical simulations. Recently, machine-learning techniques have emerged as promising tools in this…

统计力学 · 物理学 2025-02-19 Burak Çivitcioğlu , Rudolf A. Römer , Andreas Honecker
‹ 上一页 1 8 9 10 下一页 ›