English
Related papers

Related papers: An Adiabatic Theorem for Policy Tracking with TD-l…

200 papers

Temporal difference learning and Residual Gradient methods are the most widely used temporal difference based learning algorithms; however, it has been shown that none of their objective functions is optimal w.r.t approximating the true…

Machine Learning · Computer Science 2017-04-21 Bo Liu , Daoming Lyu , Wen Dong , Saad Biaz

In this paper we extend temporal difference policy evaluation algorithms to performance criteria that include the variance of the cumulative reward. Such criteria are useful for risk management, and are important in domains such as finance…

Machine Learning · Computer Science 2013-10-15 Aviv Tamar , Dotan Di Castro , Shie Mannor

Temporal difference (TD) learning is a fundamental technique in reinforcement learning that updates value estimates for states or state-action pairs using a TD target. This target represents an improved estimate of the true value by…

Machine Learning · Computer Science 2024-08-05 Wuhao Wang , Zhiyong Chen , Lepeng Zhang

The adiabatic theorem provides sufficient conditions for the time needed to prepare a target ground state. While it is possible to prepare a target state much faster with more general quantum annealing protocols, rigorous results beyond the…

Quantum Physics · Physics 2023-11-28 Luis Pedro García-Pintos , Lucas T. Brady , Jacob Bringewatt , Yi-Kai Liu

Learning the value function of a given policy from data samples is an important problem in Reinforcement Learning. TD($\lambda$) is a popular class of algorithms to solve this problem. However, the weights assigned to different $n$-step…

Machine Learning · Computer Science 2021-11-24 Rohan Deb , Meet Gandhi , Shalabh Bhatnagar

In this thesis, it is presented a set of results in adiabatic dynamics (closed and open system) and transitionless quantum driving that promote some advances in our understanding on quantum control and Hamiltonian inverse engineering. In…

Quantum Physics · Physics 2021-07-27 Alan C. Santos

The accelerated optical lattice has emerged as a valuable technique for the investigation of quantum transport physics and has found widespread application in quantum sensing, including atomic gravimeters and atomic gyroscopes. In our…

Quantum Gases · Physics 2023-09-14 Guoling Yin , Lingchii Kong , Zhongcheng Yu , Jinyuan Tian , Xuzong Chen , Xiaoji Zhou

Animals learn the timing between consecutive events very easily. Their precision is usually proportional to the interval to time (Weber's law for timing). Most current timing models either require a central clock and unbounded accumulator…

Neurons and Cognition · Quantitative Biology 2011-03-15 Francois Rivest , Yoshua Bengio

Most present applications of time-dependent density functional theory use adiabatic functionals, i.e. the effective potential at time t is determined solely by the density at the same time. This paper discusses a method that aims to go…

Strongly Correlated Electrons · Physics 2009-11-10 Yair Kurzweil , Roi Baer

We show that recent results on adiabatic theory for interacting gapped many-body systems on finite lattices remain valid in the thermodynamic limit. More precisely, we prove a generalised super-adiabatic theorem for the automorphism group…

Mathematical Physics · Physics 2024-06-19 Joscha Henheik , Stefan Teufel

Long-horizon tasks, which have a large discount factor, pose a challenge for most conventional reinforcement learning (RL) algorithms. Algorithms such as Value Iteration and Temporal Difference (TD) learning have a slow convergence rate and…

Machine Learning · Computer Science 2024-09-04 Mark Bedaywi , Amin Rakhsha , Amir-massoud Farahmand

This paper is concerned with the problem of policy evaluation with linear function approximation in discounted infinite horizon Markov decision processes. We investigate the sample complexities required to guarantee a predefined estimation…

Machine Learning · Statistics 2024-05-03 Gen Li , Weichen Wu , Yuejie Chi , Cong Ma , Alessandro Rinaldo , Yuting Wei

By introducing a temporal change timescale $\tau_{\text{A}}(t)$ for the time-dependent system Hamiltonian, a general formulation of the Markovian quantum master equation is given to go well beyond the adiabatic regime. In appropriate…

Statistical Mechanics · Physics 2017-02-01 Makoto Yamaguchi , Tatsuro Yuge , Tetsuo Ogawa

Temporal credit assignment in reinforcement learning is challenging due to delayed and stochastic outcomes. Monte Carlo targets can bridge long delays between action and consequence but lead to high-variance targets due to stochasticity.…

Machine Learning · Computer Science 2024-06-05 Aditya A. Ramesh , Kenny Young , Louis Kirsch , Jürgen Schmidhuber

Temporal logic rules are often used in control and robotics to provide structured, human-interpretable descriptions of trajectory data. These rules have numerous applications including safety validation using formal methods, constraining…

Machine Learning · Computer Science 2025-04-29 Emi Soroka , Rohan Sinha , Sanjay Lall

We establish adiabatic theorems with and without spectral gap condition for general -- typically dissipative -- linear operators $A(t): D(A(t)) \subset X \to X$ with time-dependent domains $D(A(t))$ in some Banach space $X$. In these…

Mathematical Physics · Physics 2018-09-18 Jochen Schmid

We derive a family of risk-sensitive reinforcement learning methods for agents, who face sequential decision-making tasks in uncertain environments. By applying a utility function to the temporal difference (TD) error, nonlinear…

Machine Learning · Computer Science 2014-10-10 Yun Shen , Michael J. Tobia , Tobias Sommer , Klaus Obermayer

Data-driven model predictive control has two key advantages over model-free methods: a potential for improved sample efficiency through model learning, and better performance as computational budget for planning increases. However, it is…

Machine Learning · Computer Science 2022-07-21 Nicklas Hansen , Xiaolong Wang , Hao Su

The adiabatic theorem states that an initial eigenstate of a slowly varying Hamiltonian remains close to an instantaneous eigenstate of the Hamiltonian at a later time. We show that a perfunctory application of this statement is problematic…

Quantum Physics · Physics 2009-11-10 Karl-Peter Marzlin , Barry C. Sanders

Temporal-Difference (TD) learning methods, such as Q-Learning, have proven effective at learning a policy to perform control tasks. One issue with methods like Q-Learning is that the value update introduces bias when predicting the TD…

Machine Learning · Computer Science 2021-10-29 Litian Liang , Yaosheng Xu , Stephen McAleer , Dailin Hu , Alexander Ihler , Pieter Abbeel , Roy Fox