English
Related papers

Related papers: Generalized Emphatic Temporal Difference Learning:…

200 papers

We introduce a new benchmark problem called Deceptive Leading Blocks (DLB) to rigorously study the runtime of the Univariate Marginal Distribution Algorithm (UMDA) in the presence of epistasis and deception. We show that simple Evolutionary…

Neural and Evolutionary Computing · Computer Science 2019-07-30 Per Kristian Lehre , Phan Trung Hai Nguyen

We extend the options framework for temporal abstraction in reinforcement learning from discounted Markov decision processes (MDPs) to average-reward MDPs. Our contributions include general convergent off-policy inter-option learning…

Machine Learning · Computer Science 2021-10-27 Yi Wan , Abhishek Naik , Richard S. Sutton

This paper revisits the temporal difference (TD) learning algorithm for the policy evaluation tasks in reinforcement learning. Typically, the performance of TD(0) and TD($\lambda$) is very sensitive to the choice of stepsizes. Oftentimes,…

Optimization and Control · Mathematics 2021-10-12 Tao Sun , Han Shen , Tianyi Chen , Dongsheng Li

Reinforcement Learning Algorithms are predominantly developed for stationary environments, and the limited literature that considers nonstationary environments often involves specific assumptions about changes that can occur in transition…

Machine Learning · Computer Science 2025-09-25 Ranga Shaarad Ayyagari , Revanth Raj Eega , Ambedkar Dukkipati

In this work, we study the asymptotic randomness of an algorithmic estimator of the saddle point of a globally convex-concave and locally strongly-convex strongly-concave objective. Specifically, we show that the averaged iterates of a…

Optimization and Control · Mathematics 2023-11-07 Abhishek Roy , Yi-An Ma

We propose a general formulation of a univariate estimation-of-distribution algorithm (EDA). It naturally incorporates the three classic univariate EDAs \emph{compact genetic algorithm}, \emph{univariate marginal distribution algorithm} and…

Neural and Evolutionary Computing · Computer Science 2022-10-07 Benjamin Doerr , Marc Dufay

Reinforcement learning in partially observed Markov decision processes (POMDPs) faces two challenges. (i) It often takes the full history to predict the future, which induces a sample complexity that scales exponentially with the horizon.…

Machine Learning · Computer Science 2024-04-02 Lingxiao Wang , Qi Cai , Zhuoran Yang , Zhaoran Wang

Gradient temporal-difference methods provide stable off-policy prediction with linear function approximation, but their practical performance is strongly affected by the geometry induced by the auxiliary-variable metric. Existing…

Artificial Intelligence · Computer Science 2026-05-29 Xingguo Chen , Yuchen Shen , Shangdong Yang , Chao Li , Guang Yang , Wenhao Wang

It has been recently shown in the literature that the sample averages from online learning experiments are biased when used to estimate the mean reward. To correct the bias, off-policy evaluation methods, including importance sampling and…

Machine Learning · Computer Science 2021-12-02 Ningyuan Chen , Xuefeng Gao , Yi Xiong

Production AI systems often operate with incomplete, conflicting, or insufficient evidence. Forced classifiers collapse such cases into action labels, while generative systems can produce outputs that are difficult to interpret as auditable…

Artificial Intelligence · Computer Science 2026-05-28 Sankaranarayanan Palamadai Chandrasekaran

In this work, we consider the off-policy policy evaluation problem for contextual bandits and finite horizon reinforcement learning in the nonstationary setting. Reusing old data is critical for policy evaluation, but existing estimators…

Machine Learning · Computer Science 2023-02-24 Vincent Liu , Yash Chandak , Philip Thomas , Martha White

We study the off-policy evaluation problem---estimating the value of a target policy using data collected by another policy---under the contextual bandit model. We consider the general (agnostic) setting without access to a consistent model…

Machine Learning · Statistics 2017-11-15 Yu-Xiang Wang , Alekh Agarwal , Miroslav Dudik

The focus of this paper is on stochastic variational inequalities (VI) under Markovian noise. A prominent application of our algorithmic developments is the stochastic policy evaluation problem in reinforcement learning. Prior…

Optimization and Control · Mathematics 2021-08-17 Georgios Kotsalis , Guanghui Lan , Tianjiao Li

The paper [12] examines a concept of equilibrium policies instead of optimal controls in stochastic optimization to analyze a mean-variance portfolio selection problem. We follow the same approach in order to investigate the Merton…

Optimization and Control · Mathematics 2020-04-23 I. Alia , F. Chighoub , N. Khelfallah , J. Vives

Gradient temporal-difference (GTD) learning algorithms are widely used for off-policy policy evaluation with function approximation. However, existing convergence analyses rely on the restrictive assumption that the so-called feature…

Machine Learning · Computer Science 2026-05-11 Hyunjun Na , Donghwan Lee

The Mixture Transition Distribution (MTD) model was introduced by Raftery to face the need for parsimony in the modeling of high-order Markov chains in discrete time. The particularity of this model comes from the fact that the effect of…

Computation · Statistics 2008-12-18 Sophie Lèbre , Pierre-Yves Bourguinon

We initiate the study of federated reinforcement learning under environmental heterogeneity by considering a policy evaluation problem. Our setup involves $N$ agents interacting with environments that share the same state and action space…

Machine Learning · Computer Science 2024-07-02 Han Wang , Aritra Mitra , Hamed Hassani , George J. Pappas , James Anderson

Stochastic approximation (SA) is a key method used in statistical learning. Recently, its non-asymptotic convergence analysis has been considered in many papers. However, most of the prior analyses are made under restrictive assumptions…

Machine Learning · Statistics 2019-06-18 Belhal Karimi , Blazej Miasojedow , Eric Moulines , Hoi-To Wai

Value function approximation is a crucial module for policy evaluation in reinforcement learning when the state space is large or continuous. The present paper takes a generative perspective on policy evaluation via temporal-difference (TD)…

Machine Learning · Statistics 2021-12-03 Qin Lu , Georgios B. Giannakis

A methodology of adaptive time series analysis based on Empirical Mode Decomposition (EMD) has been employed to investigate $^{7}$Be activity concentration variability, along with temperature. Analysed data were sampled at ground level by…

Geophysics · Physics 2019-05-22 Alessandro Longo , Stefano Bianchi , Wolfango Plastino
‹ Prev 1 4 5 6 7 8 10 Next ›