Related papers: An Adiabatic Theorem for Policy Tracking with TD-l…
A potential problem with adiabatic switching in perturbation theory is that divergent terms appear in the series solution. An example of this was presented by C. Brouder et al [4] for a simple 2 state system where the evolution of system in…
This paper continues the discussion on the stability of time-inhomogeneous Markov chains. In particular, this paper defines a time-inhomogeneous, discrete-time Markov chain governed by a continuous evolution in the appropriate martrix…
Off-policy learning ability is an important feature of reinforcement learning (RL) for practical applications. However, even one of the most elementary RL algorithms, temporal-difference (TD) learning, is known to suffer form divergence…
We consider a time dependent trap externally manipulated in such a way that one of its bound states is brought into an instant contact with the continuum threshold, and then down again. It is shown that, in the limit of slow evolution, the…
The hallmark feature of temporal-difference (TD) learning is bootstrapping: using value predictions to generate new value predictions. The vast majority of TD methods for control learn a policy by bootstrapping from a single action-value…
The quantum adiabatic theorem is fundamental to time dependent quantum systems, but being able to characterize quantitatively an adiabatic evolution in many-body systems can be a challenge. This work demonstrates that the use of appropriate…
The task of predicting long-term patient outcomes using supervised machine learning is a challenging one, in part because of the high variance of each patient's trajectory, which can result in the model over-fitting to the training data.…
Motivated by the emerging use of multi-agent reinforcement learning (MARL) in engineering applications such as networked robotics, swarming drones, and sensor networks, we investigate the policy evaluation problem in a fully decentralized…
We formulate an adiabatic theorem adapted to models that present an instantaneous eigenvalue experiencing an infinite number of crossings with the rest of the spectrum. We give an upper bound on the leading correction terms with respect to…
We provide rigorous bounds for the error of the adiabatic approximation of quantum mechanics under four sources of experimental error: perturbations in the initial condition, systematic time-dependent perturbations in the Hamiltonian,…
In adiabatic quantum computing finding the dependence of the gap of the Hamiltonian as a function of the parameter varied during the adiabatic sweep is crucial in order to optimize the speed of the computation. Inspired by this challenge,…
We study the convergence behavior of the celebrated temporal-difference (TD) learning algorithm. By looking at the algorithm through the lens of optimization, we first argue that TD can be viewed as an iterative optimization algorithm where…
We consider quantum field theoretic systems subject to a time-dependent perturbation, and discuss the question of defining a time dependent particle number not just at asymptotic early and late times, but also during the perturbation.…
While tabular machine learning has achieved remarkable success, temporal distribution shifts pose significant challenges in real-world deployment, as the relationships between features and labels continuously evolve. Static models assume…
We derive a Markovian master equation that models the evolution of systems subject to driving and control fields. Our approach combines time rescaling and weak-coupling limits for the system-environment interaction with a secular…
Temporal difference (TD) learning is an important approach in reinforcement learning, as it combines ideas from dynamic programming and Monte Carlo methods in a way that allows for online and incremental model-free learning. A key idea of…
We investigate whether Jacobi preconditioning, accounting for the bootstrap term in temporal difference (TD) learning, can help boost performance of adaptive optimizers. Our method, TDprop, computes a per parameter learning rate based on…
Apprenticeship learning crucially depends on effectively learning rewards, and hence control policies from user demonstrations. Of particular difficulty is the setting where the desired task consists of a number of sub-goals with temporal…
Off-policy algorithms, in which a behavior policy differs from the target policy and is used to gain experience for learning, have proven to be of great practical value in reinforcement learning. However, even for simple convex problems…
We present an analysis of the adiabatic approximation to understand when it applies, in view of the recent criticisms and studies for the validity of the adiabatic theorem. We point out that this approximation is just the leading order of a…