English
Related papers

Related papers: Exploring TD error as a heuristic for $\sigma$ sel…

200 papers

In real-world optimisation, it is common to face several sub-problems interacting and forming the main problem. There is an inter-dependency between the sub-problems, making it impossible to solve such a problem by focusing on only one…

Neural and Evolutionary Computing · Computer Science 2023-01-05 Adel Nikfarjam , Aneta Neumann , Frank Neumann

We consider LSTD($\lambda$), the least-squares temporal-difference algorithm with eligibility traces algorithm proposed by Boyan (2002). It computes a linear approximation of the value function of a fixed policy in a large Markov Decision…

Machine Learning · Computer Science 2014-05-14 Manel Tagorti , Bruno Scherrer

We develop a form Thompson sampling for online learning under full feedback - also known as prediction with expert advice - where the learner's prior is defined over the space of an adversary's future actions, rather than the space of…

Machine Learning · Computer Science 2025-09-23 Alexander Terenin , Jeffrey Negrea

The propagation of errors severely compromises the reliability of quantum computations. The quantum adiabatic algorithm is a physically motivated method to prepare ground states of classical and quantum Hamiltonians. Here, we analyze the…

Quantum Physics · Physics 2024-04-25 Benjamin F. Schiffer , Adrian Franco Rubio , Rahul Trivedi , J. Ignacio Cirac

Multi-step methods such as Retrace($\lambda$) and $n$-step $Q$-learning have become a crucial component of modern deep reinforcement learning agents. These methods are often evaluated as a part of bigger architectures and their evaluations…

Machine Learning · Computer Science 2019-02-11 J. Fernando Hernandez-Garcia , Richard S. Sutton

We introduce a variant of Multicut Decomposition Algorithms (MuDA), called CuSMuDA (Cut Selection for Multicut Decomposition Algorithms), for solving multistage stochastic linear programs that incorporates a class of cut selection…

Optimization and Control · Mathematics 2019-07-23 Vincent Guigues , Michelle Bandarra

Many reinforcement learning approaches rely on temporal-difference (TD) learning to learn a critic. However, TD-learning updates can be high variance. Here, we introduce a model-based RL framework, Taylor TD, which reduces this variance in…

Machine Learning · Computer Science 2023-10-19 Michele Garibbo , Maxime Robeyns , Laurence Aitchison

We describe an algorithm for quantum state tomography that converges in polynomial time to an estimate, together with a rigorous error bound on the fidelity between the estimate and the true state. The result suggests that state tomography…

Quantum Physics · Physics 2010-02-23 Steven T. Flammia , David Gross , Stephen D. Bartlett , Rolando Somma

In this note, we present a version of the Thompson sampling algorithm for the problem of online linear generalization with full information (i.e., the experts setting), studied by Kalai and Vempala, 2005. The algorithm uses a Gaussian prior…

Machine Learning · Statistics 2013-11-05 Aditya Gopalan

Temporal-Difference (TD) learning is a general and very useful tool for estimating the value function of a given policy, which in turn is required to find good policies. Generally speaking, TD learning updates states whenever they are…

Machine Learning · Computer Science 2021-08-24 Nishanth Anand , Doina Precup

Quantum phase estimation is one of the key algorithms in the field of quantum computing, but up until now, only approximate expressions have been derived for the probability of error. We revisit these derivations, and find that by ensuring…

Quantum Physics · Physics 2012-02-13 James M. Chappell , Max A. Lohe , Lorenz von Smekal , Azhar Iqbal , Derek Abbott

Consider $K$ processes, each generating a sequence of identical and independent random variables. The probability measures of these processes have random parameters that must be estimated. Specifically, they share a parameter $\theta$…

Machine Learning · Computer Science 2022-10-12 Arpan Mukherjee , Ali Tajer , Pin-Yu Chen , Payel Das

Motivated by a fundamental paradigm in cryptography, we consider a recent variant of the classic problem of bounding the distinguishing advantage between a random function and a random permutation. Specifically, we consider the problem of…

Information Theory · Computer Science 2020-04-22 Ido Shahaf , Or Ordentlich , Gil Segev

Consider the problem of finding a population or a probability distribution amongst many with the largest mean when these means are unknown but population samples can be simulated or otherwise generated. Typically, by selecting largest…

Probability · Mathematics 2018-09-11 Peter Glynn , Sandeep Juneja

We study on-line strategies for solving problems with hybrid algorithms. There is a problem Q and w basic algorithms for solving Q. For some lambda <= w, we have a computer with lambda disjoint memory areas, each of which can be used to run…

Discrete Mathematics · Computer Science 2007-05-23 Ming-Yang Kao , Yuan Ma , Michael Sipser , Yiqun Yin

Time-inhomogeneous finite-horizon Markov decision processes (MDP) are frequently employed to model decision-making in dynamic treatment regimes and other statistical reinforcement learning (RL) scenarios. These fields, especially healthcare…

Machine Learning · Computer Science 2025-10-21 Elynn Chen , Sai Li , Michael I. Jordan

We show that the algorithm to extract diverse M -solutions from a Conditional Random Field (called divMbest [1]) takes exactly the form of a Herding procedure [2], i.e. a deterministic dynamical system that produces a sequence of hypotheses…

Computer Vision and Pattern Recognition · Computer Science 2017-01-31 Ece Ozkan , Gemma Roig , Orcun Goksel , Xavier Boix

Temporal difference (TD) learning is a fundamental algorithm for estimating value functions in reinforcement learning. Recent finite-time analyses of TD with linear function approximation quantify its theoretical convergence rate. However,…

Machine Learning · Computer Science 2026-03-04 Yunxiang Li , Mark Schmidt , Reza Babanezhad , Sharan Vaswani

``Sim2real gap", in which the system learned in simulations is not the exact representation of the real system, can lead to loss of stability and performance when controllers learned using data from the simulated system are used on the real…

Systems and Control · Electrical Eng. & Systems 2025-05-15 Shivam Bajaj , Prateek Jaiswal , Vijay Gupta

Quantum Hamiltonian identification is important for characterizing the dynamics of quantum systems, calibrating quantum devices and achieving precise quantum control. In this paper, an effective two-step optimization (TSO) quantum…

Quantum Physics · Physics 2018-06-05 Yuanlong Wang , Daoyi Dong , Bo Qi , Jun Zhang , Ian R. Petersen , Hidehiro Yonezawa
‹ Prev 1 3 4 5 6 7 10 Next ›