English
Related papers

Related papers: An Adiabatic Theorem for Policy Tracking with TD-l…

200 papers

A general quantum adiabatic theorem with and without the time-dependent orthogonalization is proven, which can be applied to understand the origin of activation energies in chemical reactions. Further proofs are also developed for the…

Strongly Correlated Electrons · Physics 2011-11-03 Andrew Das Arulsamy

Adiabatic quantum computation provides an alternative approach to quantum computation using a time-dependent Hamiltonian. The time evolution of entanglement during the adiabatic quantum search algorithm is studied, and its relevance as a…

Quantum Physics · Physics 2009-11-11 Daria Ahrensmeier

Many facts come with an expiration date, from the name of the President to the basketball team Lebron James plays for. But language models (LMs) are trained on snapshots of data collected at a specific moment in time, and this can limit…

Computation and Language · Computer Science 2022-04-26 Bhuwan Dhingra , Jeremy R. Cole , Julian Martin Eisenschlos , Daniel Gillick , Jacob Eisenstein , William W. Cohen

This work provides test error bounds for iterative fixed point methods on linear predictors -- specifically, stochastic and batch mirror descent (MD), and stochastic temporal difference learning (TD) -- with two core contributions: (a) a…

Machine Learning · Computer Science 2022-06-29 Matus Telgarsky

Temporal difference (TD) learning is a cornerstone of reinforcement learning. In the average-reward setting, standard TD($\lambda$) is highly sensitive to the choice of step-size and thus requires careful tuning to maintain numerical…

Machine Learning · Statistics 2025-10-08 Hwanwoo Kim , Dongkyu Derek Cho , Eric Laber

We analyse quantile temporal-difference learning (QTD), a distributional reinforcement learning algorithm that has proven to be a key component in several successful large-scale applications of reinforcement learning. Despite these…

In this paper, we study the dynamics of temporal difference learning with neural network-based value function approximation over a general state space, namely, \emph{Neural TD learning}. We consider two practically used algorithms,…

Machine Learning · Computer Science 2021-08-09 Semih Cayci , Siddhartha Satpathi , Niao He , R. Srikant

Suppose an online platform wants to compare a treatment and control policy, e.g., two different matching algorithms in a ridesharing system, or two different inventory management algorithms in an online retail site. Standard randomized…

Methodology · Statistics 2022-12-27 Peter Glynn , Ramesh Johari , Mohammad Rasouli

One of the difficulties in adiabatic quantum computation is the limit on the computation time. Here we propose two schemes to speed-up the adiabatic evolution. To apply this controlled adiabatic evolution to adiabatic quantum computation,…

Quantum Physics · Physics 2015-05-14 W. Wang , S. C. Hou , X. X. Yi

One of the main obstacles to broad application of reinforcement learning methods is the parameter sensitivity of our core learning algorithms. In many large-scale applications, online computation and function approximation represent key…

Artificial Intelligence · Computer Science 2016-10-25 Martha White , Adam White

In multivariable time series (MTS) forecasting, existing state-of-the-art deep learning approaches tend to focus on autoregressive formulations and often overlook the potential of using exogenous variables in enhancing the prediction of the…

Machine Learning · Computer Science 2025-04-03 Yuxuan Shu , Vasileios Lampos

We prove an adiabatic theorem for the ground state of the Dicke model in a slowly rotating magnetic field and show that for weak electron-photon coupling, the adiabatic time scale is close to the time scale of the corresponding two level…

Functional Analysis · Mathematics 2009-10-31 J. E. Avron , A. Elgart

We apply adiabatic theorems developed for quantum mechanics to stochastic annealing processes described by the classical master equation with a time-dependent generator. When the instantaneous stationary state is unique and the minimum…

Statistical Mechanics · Physics 2024-03-21 Kazutaka Takahashi

Modern learning systems increasingly interact with data that evolve over time and depend on hidden internal state. We ask a basic question: when is such a dynamical system learnable from observations alone? This paper proposes a research…

Machine Learning · Computer Science 2025-12-23 Elad Hazan , Shai Shalev Shwartz , Nathan Srebro

Temporal difference learning (TD) is a foundational concept in reinforcement learning (RL), aimed at efficiently assessing a policy's value function. TD($\lambda$), a potent variant, incorporates a memory trace to distribute the prediction…

Machine Learning · Computer Science 2024-02-13 Jianfei Ma

The condition for adiabatic approximation are of basic importance for the applications of the adiabatic theorem. The traditional quantitative condition was found to be necessary but not sufficient, but we do not know its physical meaning…

Quantum Physics · Physics 2011-02-02 Qian-Heng Duan , Ping-Xing Chen , Wei Wu

Adiabatic quantum control protocols have been of wide interest to quantum computation due to their robustness and insensitivity to their actual duration of execution. As an extension of previous quantum learning algorithms, this work…

Quantum Physics · Physics 2023-03-03 Nannan Ma , Wenhao Chu , Jiangbin Gong

We study the policy evaluation problem in multi-agent reinforcement learning. In this problem, a group of agents works cooperatively to evaluate the value function for the global discounted accumulative reward problem, which is composed of…

Optimization and Control · Mathematics 2019-06-04 Thinh T. Doan , Siva Theja Maguluri , Justin Romberg

Autonomous agents operating in continuous environments must decide not only what to do, but when to act. We introduce a lightweight adaptive temporal control system that learns the optimal interval between cognitive ticks from experience,…

Machine Learning · Computer Science 2026-03-27 Davide Di Gioia

Consider the problem of learning the drift coefficient of a stochastic differential equation from a sample path. In this paper, we assume that the drift is parametrized by a high dimensional vector. We address the question of how long the…

Information Theory · Computer Science 2011-03-10 José Bento , Morteza Ibrahimi , Andrea Montanari
‹ Prev 1 4 5 6 7 8 10 Next ›