中文
相关论文

相关论文: Krylov-Bellman boosting: Super-linear policy evalu…

200 篇论文

We study reinforcement learning methods with linear function approximation under non-Markov state and cost processes. We first consider the policy evaluation method and show that the algorithm converges under suitable ergodicity conditions…

机器学习 · 计算机科学 2026-01-05 Ali Devran Kara

The Krylov subspace method is a standard approach to approximate quantum evolution, allowing to treat systems with large Hilbert spaces. Although its application is general, and suitable for many-body systems, estimation of the committed…

量子物理 · 物理学 2021-07-22 Julian Ruffinelli , Emiliano Fortes , Martín Larocca , Diego A. Wisniacki

In recent years, reinforcement learning (RL) systems with general goals beyond a cumulative sum of rewards have gained traction, such as in constrained problems, exploration, and acting upon prior experiences. In this paper, we consider…

机器学习 · 计算机科学 2020-07-07 Junyu Zhang , Alec Koppel , Amrit Singh Bedi , Csaba Szepesvari , Mengdi Wang

We describe a nonlinear generalization of dual dynamic programming theory and its application to value function estimation for deterministic control problems over continuous state and action spaces, in a discrete-time infinite horizon…

最优化与控制 · 数学 2018-10-05 Joseph Warrington , Paul N. Beuchat , John Lygeros

We prove that boosting with the squared error loss, $L_2$Boosting, is consistent for very high-dimensional linear models, where the number of predictor variables is allowed to grow essentially as fast as $O$(exp(sample size)), assuming that…

统计理论 · 数学 2016-08-16 Peter Bühlmann

Policy evaluation is a key process in reinforcement learning. It assesses a given policy using estimation of the corresponding value function. When using a parameterized function to approximate the value, it is common to optimize the set of…

机器学习 · 计算机科学 2019-01-24 Shirli Di-Castro Shashua , Shie Mannor

We address the problem of automatic generation of features for value function approximation. Bellman Error Basis Functions (BEBFs) have been shown to improve the error of policy evaluation with function approximation, with a convergence…

机器学习 · 计算机科学 2012-09-25 Mahdi Milani Fard , Yuri Grinberg , Amir-massoud Farahmand , Joelle Pineau , Doina Precup

We present the first finite-sample analysis of policy evaluation in robust average-reward Markov Decision Processes (MDPs). Prior work in this setting have established only asymptotic convergence guarantees, leaving open the question of…

机器学习 · 统计学 2025-12-11 Yang Xu , Washim Uddin Mondal , Vaneet Aggarwal

Boosting is one of the most significant developments in machine learning. This paper studies the rate of convergence of $L_2$Boosting, which is tailored for regression, in a high-dimensional setting. Moreover, we introduce so-called…

机器学习 · 统计学 2022-07-22 Ye Luo , Martin Spindler , Jannis Kück

Boosting algorithms to simultaneously estimate and select predictor effects in statistical models have gained substantial interest during the last decade. This review article aims to highlight recent methodological developments regarding…

统计方法学 · 统计学 2014-11-19 Andreas Mayr , Harald Binder , Olaf Gefeller , Matthias Schmid

Quantum Krylov subspace diagonalization (QKSD) is an emerging method used in place of quantum phase estimation in the early fault-tolerant era, where limited quantum circuit depth is available. In contrast to the classical Krylov subspace…

量子物理 · 物理学 2024-09-20 Gwonhak Lee , Dongkeun Lee , Joonsuk Huh

We introduce a practical method for incorporating equality and inequality constraints in global optimization methods based on stochastic interacting particle systems, specifically consensus-based optimization (CBO) and ensemble Kalman…

最优化与控制 · 数学 2021-11-05 J. A. Carrillo , C. Totzeck , U. Vaes

In this paper, we study a Markov decision process with a non-linear discount function and with a Borel state space. We define a recursive discounted utility, which resembles non-additive utility functions considered in a number of models in…

最优化与控制 · 数学 2025-10-16 Nicole Bäuerle , Anna Jaśkiewicz , Andrzej S. Nowak

We seek to learn an effective policy for a Markov Decision Process (MDP) with continuous states via Q-Learning. Given a set of basis functions over state action pairs we search for a corresponding set of linear weights that minimizes the…

机器学习 · 计算机科学 2013-09-27 Charles Tripp , Ross D. Shachter

We propose a refinement of temporal-difference learning that enforces first-order Bellman consistency: the learned value function is trained to match not only the Bellman targets in value but also their derivatives with respect to states…

机器学习 · 计算机科学 2025-11-25 Fabian Schramm , Nicolas Perrin-Gilbert , Justin Carpentier

Prediction error is critical to assessing the performance of statistical methods and selecting statistical models. We propose the cross-validation and approximated cross-validation methods for estimating prediction error under a broad…

统计理论 · 数学 2007-06-13 Jianqing Fan , Chunming Zhang

We study the estimation of the value function for continuous-time Markov diffusion processes using a single, discretely observed ergodic trajectory. Our work provides non-asymptotic statistical guarantees for the least-squares…

机器学习 · 计算机科学 2025-02-07 Wenlong Mou

In this paper we analyze boosting algorithms in linear regression from a new perspective: that of modern first-order methods in convex optimization. We show that classic boosting algorithms in linear regression, namely the incremental…

统计理论 · 数学 2015-05-19 Robert M. Freund , Paul Grigas , Rahul Mazumder

We propose a formulation of the stochastic cutting stock problem as a discounted infinite-horizon Markov decision process. At each decision epoch, given current inventory of items, an agent chooses in which patterns to cut objects in stock…

最优化与控制 · 数学 2022-06-29 Anselmo R. Pitombeira-Neto , Arthur H. Fonseca Murta

Continuous-time Markov decision processes are an important class of models in a wide range of applications, ranging from cyber-physical systems to synthetic biology. A central problem is how to devise a policy to control the system in order…

系统与控制 · 计算机科学 2016-06-01 Ezio Bartocci , Luca Bortolussi , Tomǎš Brázdil , Dimitrios Milios , Guido Sanguinetti