Related papers: An Adiabatic Theorem for Policy Tracking with TD-l…
Adiabatic approximations break down classically when a constant-energy contour splits into separate contours, forcing the system to choose which daughter contour to follow; the choices often represent qualitatively different behavior, so…
Temporal-difference learning (TD), coupled with neural networks, is among the most fundamental building blocks of deep reinforcement learning. However, due to the nonlinearity in value function approximation, such a coupling leads to…
We revisit the time-adiabatic theorem of quantum mechanics and show that it can be extended to weakly nonlinear situations, that is to nonlinear Schroedinger equations in which either the nonlinear coupling constant or, equivalently, the…
We show that the time dependent single electron, nuclear density matrix of an interacting electronic system coupled to nuclear degrees of freedom can be exactly reproduced by that of an electronic system with arbitrarily specified…
In continuing tasks, average-reward reinforcement learning may be a more appropriate problem formulation than the more common discounted reward formulation. As usual, learning an optimal policy in this setting typically requires a large…
This paper investigates the problem of online prediction learning, where learning proceeds continuously as the agent interacts with an environment. The predictions made by the agent are contingent on a particular way of behaving,…
Dynamic nonlinear systems exhibit distortions arising from coupled static and dynamic effects. Their intertwined nature poses major challenges for data-driven modeling. This paper presents a theoretical framework grounded in structured…
We review the quantum adiabatic approximation for closed systems, and its recently introduced generalization to open systems (M.S. Sarandy and D.A. Lidar, e-print quant-ph/0404147). We also critically examine a recent argument claiming that…
We prove a non-asymptotic central limit theorem for vector-valued martingale differences using Stein's method, and use Poisson's equation to extend the result to functions of Markov Chains. We then show that these results can be applied to…
Given the rapidly evolving nature of social media and people's views, word usage changes over time. Consequently, the performance of a classifier trained on old textual data can drop dramatically when tested on newer data. While research in…
A novel way of using neural networks to learn the dynamics of time delay systems from sequential data is proposed. A neural network with trainable delays is used to approximate the right hand side of a delay differential equation. We relate…
We propose a particularly structured Boltzmann machine, which we refer to as a dynamic Boltzmann machine (DyBM), as a stochastic model of a multi-dimensional time-series. The DyBM can have infinitely many layers of units but allows exact…
The quantum speed limit specifies a universal bound of the fidelity between the initial state and the time-evolved state. We apply this method to find a bound of the fidelity between the adiabatic state and the time-evolved state. The bound…
The plasticity of the conduction delay between neurons plays a fundamental role in learning. However, the exact underlying mechanisms in the brain for this modulation is still an open problem. Understanding the precise adjustment of…
We consider reinforcement learning with performance evaluated by a dynamic risk measure. We construct a projected risk-averse dynamic programming equation and study its properties. Then we propose risk-averse counterparts of the methods of…
Using insight from numerical approximation of ODEs and the problem formulation and solution methodology of TD learning through a Galerkin relaxation, I propose a new class of TD learning algorithms. After applying the improved numerical…
Tracking the solution of time-varying variational inequalities is an important problem with applications in game theory, optimization, and machine learning. Existing work considers time-varying games or time-varying optimization problems.…
This paper proposes a reinforcement learning method for controller synthesis of autonomous systems in unknown and partially-observable environments with subjective time-dependent safety constraints. Mathematically, we model the system…
In many finite horizon episodic reinforcement learning (RL) settings, it is desirable to optimize for the undiscounted return - in settings like Atari, for instance, the goal is to collect the most points while staying alive in the long…
It has been recently reported that classical systems have speed limit for state evolution, although such a concept of speed limit had been considered to be unique to quantum systems. Owing to the speed limit for classical system, the lower…