English
Related papers

Related papers: Beyond the Bellman Fixed Point: Geometry and Fast …

200 papers

We propose a novel group of Gaussian Process based algorithms for fast approximate optimal stopping of time series with specific applications to financial markets. We show that structural properties commonly exhibited by financial time…

Machine Learning · Statistics 2022-10-11 Kshama Dwarakanath , Danial Dervovic , Peyman Tavallali , Svitlana S Vyetrenko , Tucker Balch

In this paper, we study the offline RL problem with linear function approximation. Our main structural assumption is that the MDP has low inherent Bellman error, which stipulates that linear value functions have linear Bellman backups with…

Machine Learning · Computer Science 2024-06-19 Noah Golowich , Ankur Moitra

We consider the \emph{Budgeted} version of the classical \emph{Connected Dominating Set} problem (BCDS). Given a graph $G$ and a budget $k$, we seek a connected subset of at most $k$ vertices maximizing the number of dominated vertices in…

Data Structures and Algorithms · Computer Science 2020-03-26 Ioannis Lamprou , Ioannis Sigalas , Vassilis Zissimopoulos

Average-reward reinforcement learning requires estimating the gain and the bias, which is defined only up to an additive constant. This makes direct distributional analogues ill-posed on the real line. We introduce a quotient-space…

Machine Learning · Computer Science 2026-05-13 Ege C. Kaya , Aliasghar Pourghani , Vijay Gupta , Abolfazl Hashemi

We study the convergence of $Q$-learning with linear function approximation. Our key contribution is the introduction of a novel multi-Bellman operator that extends the traditional Bellman operator. By exploring the properties of this…

Machine Learning · Computer Science 2023-10-02 Diogo S. Carvalho , Pedro A. Santos , Francisco S. Melo

Q-learning is a fundamental algorithmic primitive in reinforcement learning. This paper develops a new framework for analyzing Q-learning from a switching-system viewpoint. In particular, we derive a direct stochastic switching-system…

Machine Learning · Computer Science 2026-05-06 Donghwan Lee

Given a finite family of functions, the goal of model selection aggregation is to construct a procedure that mimics the function from this family that is the closest to an unknown regression function. More precisely, we consider a general…

Statistics Theory · Mathematics 2012-12-13 Dong Dai , Philippe Rigollet , Tong Zhang

Recent policy optimization approaches have achieved substantial empirical success by constructing surrogate optimization objectives. The Approximate Policy Iteration objective (Schulman et al., 2015a; Kakade and Langford, 2002) has become a…

Machine Learning · Computer Science 2019-10-31 Marcin B. Tomczak , Sergio Valcarcel Macua , Enrique Munoz de Cote , Peter Vrancx

While off-policy temporal difference (TD) methods have widely been used in reinforcement learning due to their efficiency and simple implementation, their Bayesian counterparts have not been utilized as frequently. One reason is that the…

Machine Learning · Computer Science 2019-10-25 Heejin Jeong , Clark Zhang , George J. Pappas , Daniel D. Lee

Temporal-difference learning (TD), coupled with neural networks, is among the most fundamental building blocks of deep reinforcement learning. However, due to the nonlinearity in value function approximation, such a coupling leads to…

Machine Learning · Computer Science 2020-04-16 Qi Cai , Zhuoran Yang , Jason D. Lee , Zhaoran Wang

This paper studies the statistical theory of batch data reinforcement learning with function approximation. Consider the off-policy evaluation problem, which is to estimate the cumulative value of a new target policy from logged history…

Machine Learning · Computer Science 2020-02-25 Yaqi Duan , Mengdi Wang

This note provides upper bounds on the number of operations required to compute by value iterations a nearly optimal policy for an infinite-horizon discounted Markov decision process with a finite number of states and actions. For a given…

Optimization and Control · Mathematics 2020-01-29 Eugene A. Feinberg , Gaojin He

We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework…

Optimization and Control · Mathematics 2025-06-11 Qi Feng , Gu Wang

One of the main challenges in model-based reinforcement learning (RL) is to decide which aspects of the environment should be modeled. The value-equivalence (VE) principle proposes a simple answer to this question: a model should capture…

Artificial Intelligence · Computer Science 2021-12-14 Christopher Grimm , André Barreto , Gregory Farquhar , David Silver , Satinder Singh

The Whittle index policy is a heuristic that has shown remarkably good performance (with guaranteed asymptotic optimality) when applied to the class of problems known as Restless Multi-Armed Bandit Problems (RMABPs). In this paper we…

Artificial Intelligence · Computer Science 2024-06-05 Francisco Robledo Relaño , Vivek Borkar , Urtzi Ayesta , Konstantin Avrachenkov

In this article, we study optimal investment and consumption in an incomplete stochastic factor model for a power utility investor on the infinite horizon. When the state space of the stochastic factor is finite, we give a complete…

Mathematical Finance · Quantitative Finance 2025-09-12 Florian Gutekunst , Martin Herdegen , David Hobson

We develop methods for nonparametric uniform inference in cost-sensitive binary classification, a framework that encompasses maximum score estimation, predicting utility maximizing actions, and policy learning. These problems are well known…

Econometrics · Economics 2025-12-16 Nan Liu , Yanbo Liu , Yuya Sasaki , Yuanyuan Wan

In this paper, we derive a generalization of the Speedy Q-learning (SQL) algorithm that was proposed in the Reinforcement Learning (RL) literature to handle slow convergence of Watkins' Q-learning. In most RL algorithms such as Q-learning,…

Machine Learning · Computer Science 2020-02-14 Indu John , Chandramouli Kamanchi , Shalabh Bhatnagar

Although Q-learning is one of the most successful algorithms for finding the best action-value function (and thus the optimal policy) in reinforcement learning, its implementation often suffers from large overestimation of Q-function values…

Machine Learning · Computer Science 2020-10-13 Huaqing Xiong , Lin Zhao , Yingbin Liang , Wei Zhang

In this paper, for POMDPs, we provide the convergence of a Q learning algorithm for control policies using a finite history of past observations and control actions, and, consequentially, we establish near optimality of such limit Q…

Machine Learning · Computer Science 2022-10-27 Ali Devran Kara , Serdar Yuksel
‹ Prev 1 8 9 10 Next ›