中文
相关论文

相关论文: Properties of the Least Squares Temporal Differenc…

200 篇论文

Stochastic gradient descent (SGD) is a popular algorithm for minimizing objective functions that arise in machine learning. For constant step-sized SGD, the iterates form a Markov chain on a general state space. Focusing on a class of…

最优化与控制 · 数学 2025-03-26 David Shirokoff , Philip Zaleski

We study episodic reinforcement learning in non-stationary linear (a.k.a. low-rank) Markov Decision Processes (MDPs), i.e, both the reward and transition kernel are linear with respect to a given feature map and are allowed to evolve either…

机器学习 · 计算机科学 2021-12-28 Ahmed Touati , Pascal Vincent

Multi Task Learning (MTL) efficiently leverages useful information contained in multiple related tasks to help improve the generalization performance of all tasks. This article conducts a large dimensional analysis of a simple but, as we…

机器学习 · 统计学 2020-09-04 Malik Tiomoko , Romain Couillet , Hafiz Tiomoko

We investigate the problem of optimal control synthesis for Markov Decision Processes (MDPs), addressing both qualitative and quantitative objectives. Specifically, we require the system to satisfy a qualitative task specified by a Linear…

系统与控制 · 电气工程与系统科学 2025-09-19 Yu Chen , Xuanyuan Yin , Shaoyuan Li , Xiang Yin

Temporal difference learning (TD) is a foundational concept in reinforcement learning (RL), aimed at efficiently assessing a policy's value function. TD($\lambda$), a potent variant, incorporates a memory trace to distribute the prediction…

机器学习 · 计算机科学 2024-02-13 Jianfei Ma

We study computational and statistical aspects of learning Latent Markov Decision Processes (LMDPs). In this model, the learner interacts with an MDP drawn at the beginning of each epoch from an unknown mixture of MDPs. To sidestep known…

机器学习 · 计算机科学 2024-06-13 Fan Chen , Constantinos Daskalakis , Noah Golowich , Alexander Rakhlin

Least squares approximation is a technique to find an approximate solution to a system of linear equations that has no exact solution. In a typical setting, one lets $n$ be the number of constraints and $d$ be the number of variables, with…

数据结构与算法 · 计算机科学 2010-09-28 Petros Drineas , Michael W. Mahoney , S. Muthukrishnan , Tamas Sarlos

This paper aims to present a local discontinuous Galerkin (LDG) method for solving backward stochastic partial differential equations (BSPDEs) with Neumann boundary conditions. We establish the $L^2$-stability and optimal error estimates of…

数值分析 · 数学 2024-09-18 Yixiang Dai , Yunzhang Li , Jing Zhang

Gradient temporal difference (Gradient TD) algorithms are a popular class of stochastic approximation (SA) algorithms used for policy evaluation in reinforcement learning. Here, we consider Gradient TD algorithms with an additional heavy…

机器学习 · 计算机科学 2021-11-23 Rohan Deb , Shalabh Bhatnagar

In this paper, we study the distributed adaptive estimation problem of continuous-time stochastic dynamic systems over sensor networks where each agent can only communicate with its local neighbors. A distributed least squares (LS)…

系统与控制 · 电气工程与系统科学 2023-09-07 Xinghua Zhu , Zhixin Liu

A space-time-parameters structure of parametric parabolic PDEs motivates the application of tensor methods to define reduced order models (ROMs). Within a tensor-based ROM framework, the matrix SVD - a traditional dimension reduction…

数值分析 · 数学 2024-08-22 Alexander V. Mamonov , Maxim A. Olshanskii

TD($\lambda$) with function approximation has proved empirically successful for some complex reinforcement learning problems. For linear approximation, TD($\lambda$) has been shown to minimise the squared error between the approximate value…

机器学习 · 计算机科学 2025-12-24 Lex Weaver , Jonathan Baxter

The reduced-rank method exploits the distortion-variance tradeoff to yield superior solutions for classic problems in statistical signal processing such as parameter estimation and filtering. The central idea is to reduce the variance of…

信息论 · 计算机科学 2019-03-06 K. G. Nagananda , Pramod Khargonekar

We study reinforcement learning with linear function approximation where the transition probability and reward functions are linear with respect to a feature mapping $\boldsymbol{\phi}(s,a)$. Specifically, we consider the episodic…

机器学习 · 计算机科学 2023-01-31 Pihe Hu , Yu Chen , Longbo Huang

This work presents a new variation of the commonly used Least Mean Squares Algorithm (LMS) for the identification of sparse signals with an a-priori known sparsity using a hard threshold operator in every iteration. It examines some useful…

系统与控制 · 计算机科学 2016-08-04 Lampros Flokas , Petros Maragos

We propose a new discrete-time online parameter estimation algorithm that combines two different aspects, one that adds momentum, and another that includes a time-varying learning rate. It is well known that recursive least squares based…

最优化与控制 · 数学 2023-03-21 Yingnan Cui , Anuradha M. Annaswamy

Least-squares reverse time migration (LSRTM) is one of the classic seismic imaging methods to reconstruct model perturbations within a known reference medium. It can be computed in either data or image domain using different methods by…

地球物理 · 物理学 2025-12-11 Pengliang Yang , Zhengyu Ji

We propose a novel algorithm for greedy forward feature selection for regularized least-squares (RLS) regression and classification, also known as the least-squares support vector machine or ridge regression. The algorithm, which we call…

机器学习 · 统计学 2010-03-19 Tapio Pahikkala , Antti Airola , Tapio Salakoski

For a parameterized hyperbolic system $\frac{du}{dt}=f(u,s)$ the derivative of the ergodic average $\langle J \rangle = \lim_{T \to \infty}\frac{1}{T}\int_0^T J(u(t),s)$ to the parameter $s$ can be computed via the Least Squares Shadowing…

动力系统 · 数学 2017-09-13 Mario Chater , Angxiu Ni , Patrick J. Blonigan , Qiqi Wang

Markov parameters play a key role in system identification. There exists many algorithms where these parameters are estimated using least-squares in a first, pre-processing, step, including subspace identification and multi-step…

系统与控制 · 电气工程与系统科学 2024-05-08 Jiabao He , Cristian R. Rojas , Håkan Hjalmarsson