中文
相关论文

相关论文: A predefined-time first-order exact differentiator…

200 篇论文

We derive an equation for temporal difference learning from statistical principles. Specifically, we start with the variational principle and then bootstrap to produce an updating rule for discounted state value estimates. The resulting…

机器学习 · 计算机科学 2008-11-03 Marcus Hutter , Shane Legg

Reactive synthesis from high-level specifications that combine hard constraints expressed in Linear Temporal Logic LTL with soft constraints expressed by discounted-sum (DS) rewards has applications in planning and reinforcement learning.…

人工智能 · 计算机科学 2022-05-24 Suguman Bansal , Lydia Kavraki , Moshe Y. Vardi , Andrew Wells

Numerous Optimization Algorithms have a time-varying update rule thanks to, for instance, a changing step size, momentum parameter or, Hessian approximation. In this paper, we apply unrolled or automatic differentiation to a time-varying…

最优化与控制 · 数学 2024-10-28 Sheheryar Mehmood , Peter Ochs

We propose a new method for simulating certain type of time-dependent Hamiltonian $H(t) = \sum_{i=1}^m \gamma_i(t) H_i$ where $\gamma_i(t)$ (and its higher order derivatives) is bounded, computable function of time $t$, and each $H_i$ is…

量子物理 · 物理学 2024-10-21 Nhat A. Nghiem

In this paper, a second order finite difference scheme is investigated for time-dependent one-side space fractional diffusion equations with variable coefficients. The existing schemes for the equation with variable coefficients have…

数值分析 · 数学 2019-02-25 Xue-lei Lin , Pin Lyu , Michael K. Ng , Hai-Wei Sun , Seakweng Vong

The present paper is devoted to constructing L2 type difference analog of the Caputo fractional derivative. The fundamental features of this difference operator are studied and it is used to construct difference schemes generating…

数值分析 · 数学 2021-02-18 Anatoly A. Alikhanov , Chengming Huang

The ubiquity of time series data creates a strong demand for general-purpose foundation models, yet developing them for classification remains a significant challenge, largely due to the high cost of labeled data. Foundation models capable…

机器学习 · 计算机科学 2025-11-27 Chin-Chia Michael Yeh , Uday Singh Saini , Junpeng Wang , Xin Dai , Xiran Fan , Jiarui Sun , Yujie Fan , Yan Zheng

Stability of linear systems with uncertain bounded time-varying delays is studied under assumption that the nominal delay values are not equal to zero. An input-output approach to stability of such systems is known to be based on the bound…

最优化与控制 · 数学 2007-05-23 Eugenii Shustin , Emilia Fridman

Considering the concept of time-dilation, there exist some major issues with recurrent neural Architectures. Any variation in time spans between input data points causes performance attenuation in recurrent neural network architectures.…

计算机视觉与模式识别 · 计算机科学 2022-04-12 Aref Hakimzadeh , Koorush Ziarati , Mohammad Taheri

In this paper, we propose a direct parallel-in-time (PinT) algorithm for time-dependent problems with first- or second-order derivative. We use a second-order boundary value method as the time integrator that leads to a tridiagonal time…

数值分析 · 数学 2022-02-22 Jun Liu , Xiang-Sheng Wang , Shu-Lin Wu , Tao Zhou

A recurring theme in statistical learning, online learning, and beyond is that faster convergence rates are possible for problems with low noise, often quantified by the performance of the best hypothesis; such results are known as…

机器学习 · 计算机科学 2021-07-07 Dylan J. Foster , Akshay Krishnamurthy

We consider the problem of minimizing a differentiable function with locally Lipschitz continuous gradient on a stratified set and present a first-order algorithm designed to find a stationary point of that problem. Our assumptions on the…

最优化与控制 · 数学 2023-03-29 Guillaume Olikier , Kyle A. Gallivan , P. -A. Absil

In reinforcement learning, temporal difference (TD) is the most direct algorithm to learn the value function of a policy. For large or infinite state spaces, exact representations of the value function are usually not available, and it must…

机器学习 · 计算机科学 2018-05-03 Yann Ollivier

A novel strategy aimed at cooperatively differentiating a signal among multiple interacting agents is introduced, where none of the agents needs to know which agent is the leader, i.e. the one producing the signal to be differentiated.…

系统与控制 · 电气工程与系统科学 2025-02-14 Rodrigo Aldana-Lopez , David Gomez-Gutierrez , Elio Usai , Hernan Haimovich

The lifted dynamic junction tree algorithm (LDJT) efficiently answers filtering and prediction queries for probabilistic relational temporal models by building and then reusing a first-order cluster representation of a knowledge base for…

人工智能 · 计算机科学 2018-07-03 Marcel Gehrke , Tanya Braun , Ralf Möller

We establish novel and general high-dimensional concentration inequalities and Berry-Esseen bounds for vector-valued martingales induced by Markov chains. We apply these results to analyze the performance of the Temporal Difference (TD)…

机器学习 · 统计学 2026-05-22 Weichen Wu , Yuting Wei , Alessandro Rinaldo

This article studies the solutions of time-dependent differential inclusions which is motivated by their utility in the modeling of certain physical systems. The differential inclusion is described by a time-dependent set-valued mapping…

最优化与控制 · 数学 2021-07-05 Kanat Camlibel , Luigi Iannelli , Aneel Tanwani

This paper investigates estimating the variance of a temporal-difference learning agent's update target. Most reinforcement learning methods use an estimate of the value function, which captures how good it is for the agent to be in a…

人工智能 · 计算机科学 2018-02-15 Craig Sherstan , Brendan Bennett , Kenny Young , Dylan R. Ashley , Adam White , Martha White , Richard S. Sutton

Learning the value function of a given policy (target policy) from the data samples obtained from a different policy (behavior policy) is an important problem in Reinforcement Learning (RL). This problem is studied under the setting of…

机器学习 · 计算机科学 2019-11-14 Raghuram Bharadwaj Diddigi , Chandramouli Kamanchi , Shalabh Bhatnagar

In many applications of finance, biology and sociology, complex systems involve entities interacting with each other. These processes have the peculiarity of evolving over time and of comprising latent factors, which influence the system…

机器学习 · 统计学 2018-08-03 Federico Tomasi , Veronica Tozzo , Saverio Salzo , Alessandro Verri