Related papers: An Adiabatic Theorem for Policy Tracking with TD-l…
Temporal difference learning and Residual Gradient methods are the most widely used temporal difference based learning algorithms; however, it has been shown that none of their objective functions is optimal w.r.t approximating the true…
In this paper we extend temporal difference policy evaluation algorithms to performance criteria that include the variance of the cumulative reward. Such criteria are useful for risk management, and are important in domains such as finance…
Temporal difference (TD) learning is a fundamental technique in reinforcement learning that updates value estimates for states or state-action pairs using a TD target. This target represents an improved estimate of the true value by…
The adiabatic theorem provides sufficient conditions for the time needed to prepare a target ground state. While it is possible to prepare a target state much faster with more general quantum annealing protocols, rigorous results beyond the…
Learning the value function of a given policy from data samples is an important problem in Reinforcement Learning. TD($\lambda$) is a popular class of algorithms to solve this problem. However, the weights assigned to different $n$-step…
In this thesis, it is presented a set of results in adiabatic dynamics (closed and open system) and transitionless quantum driving that promote some advances in our understanding on quantum control and Hamiltonian inverse engineering. In…
The accelerated optical lattice has emerged as a valuable technique for the investigation of quantum transport physics and has found widespread application in quantum sensing, including atomic gravimeters and atomic gyroscopes. In our…
Animals learn the timing between consecutive events very easily. Their precision is usually proportional to the interval to time (Weber's law for timing). Most current timing models either require a central clock and unbounded accumulator…
Most present applications of time-dependent density functional theory use adiabatic functionals, i.e. the effective potential at time t is determined solely by the density at the same time. This paper discusses a method that aims to go…
We show that recent results on adiabatic theory for interacting gapped many-body systems on finite lattices remain valid in the thermodynamic limit. More precisely, we prove a generalised super-adiabatic theorem for the automorphism group…
Long-horizon tasks, which have a large discount factor, pose a challenge for most conventional reinforcement learning (RL) algorithms. Algorithms such as Value Iteration and Temporal Difference (TD) learning have a slow convergence rate and…
This paper is concerned with the problem of policy evaluation with linear function approximation in discounted infinite horizon Markov decision processes. We investigate the sample complexities required to guarantee a predefined estimation…
By introducing a temporal change timescale $\tau_{\text{A}}(t)$ for the time-dependent system Hamiltonian, a general formulation of the Markovian quantum master equation is given to go well beyond the adiabatic regime. In appropriate…
Temporal credit assignment in reinforcement learning is challenging due to delayed and stochastic outcomes. Monte Carlo targets can bridge long delays between action and consequence but lead to high-variance targets due to stochasticity.…
Temporal logic rules are often used in control and robotics to provide structured, human-interpretable descriptions of trajectory data. These rules have numerous applications including safety validation using formal methods, constraining…
We establish adiabatic theorems with and without spectral gap condition for general -- typically dissipative -- linear operators $A(t): D(A(t)) \subset X \to X$ with time-dependent domains $D(A(t))$ in some Banach space $X$. In these…
We derive a family of risk-sensitive reinforcement learning methods for agents, who face sequential decision-making tasks in uncertain environments. By applying a utility function to the temporal difference (TD) error, nonlinear…
Data-driven model predictive control has two key advantages over model-free methods: a potential for improved sample efficiency through model learning, and better performance as computational budget for planning increases. However, it is…
The adiabatic theorem states that an initial eigenstate of a slowly varying Hamiltonian remains close to an instantaneous eigenstate of the Hamiltonian at a later time. We show that a perfunctory application of this statement is problematic…
Temporal-Difference (TD) learning methods, such as Q-Learning, have proven effective at learning a policy to perform control tasks. One issue with methods like Q-Learning is that the value update introduces bias when predicting the TD…