English
Related papers

Related papers: Multi-Bellman operator for convergence of $Q$-lear…

200 papers

The problem of solving Markov decision processes under function approximation remains a fundamental challenge, even under linear function approximation settings. A key difficulty arises from a geometric mismatch: while the Bellman…

Machine Learning · Computer Science 2026-04-09 Hyukjun Yang , Han-Dong Lim , Donghwan Lee

We obtain a new universal approximation theorem for continuous (possibly nonlinear) operators on arbitrary Banach spaces using the Leray-Schauder mapping. Moreover, we introduce and study a method for operator learning in Banach spaces…

Numerical Analysis · Mathematics 2026-03-17 Emanuele Zappala

Convex optimization models find interesting applications, especially in signal/image processing and compressive sensing. We study some augmented convex models, which are perturbed by strongly convex functions, and propose a dual gradient…

Optimization and Control · Mathematics 2013-08-30 Hui Zhang , Lizhi Cheng , Wotao Yin

Q-learning is known as one of the fundamental reinforcement learning (RL) algorithms. Its convergence has been the focus of extensive research over the past several decades. Recently, a new finitetime error bound and analysis for Q-learning…

Systems and Control · Electrical Eng. & Systems 2024-01-17 Donghwna Lee

Quantum Machine Learning(QML) is developed by combining quantum mechanics principles with classical machine learning techniques in a hybrid framework that can give faster, exponential, more efficient power of quantum computing with the data…

Quantum Physics · Physics 2026-01-27 Pallab Biswas , Tamal Maity

Fully decentralized learning, where the global information, i.e., the actions of other agents, is inaccessible, is a fundamental challenge in cooperative multi-agent reinforcement learning. However, the convergence and optimality of most…

Machine Learning · Computer Science 2023-02-03 Jiechuan Jiang , Zongqing Lu

Operator learning refers to the application of ideas from machine learning to approximate (typically nonlinear) operators mapping between Banach spaces of functions. Such operators often arise from physical models expressed in terms of…

Machine Learning · Computer Science 2024-02-27 Nikola B. Kovachki , Samuel Lanthaler , Andrew M. Stuart

Reinforcement learning has witnessed significant advancements, particularly with the emergence of model-based approaches. Among these, $Q$-learning has proven to be a powerful algorithm in model-free settings. However, the extension of…

Machine Learning · Computer Science 2026-03-31 Han-Dong Lim , HyeAnn Lee , Donghwan Lee

In this paper, we formalize the almost sure convergence of $Q$-learning and linear temporal difference (TD) learning with Markovian samples using the Lean 4 theorem prover based on the Mathlib library. $Q$-learning and linear TD are among…

Machine Learning · Computer Science 2025-11-06 Shangtong Zhang

In this paper, we investigate the problem of controlling probabilistic Boolean control networks (PBCNs) to achieve reachability with maximum probability in the finite time horizon. We address three questions: 1) finding control policies…

Systems and Control · Electrical Eng. & Systems 2023-12-13 Hongyue Fan , Jingjie Ni , Fangfei Li

We introduce a reinforcement learning algorithm designed to identify the fixed points of a given quantum operation. The method iteratively constructs the unitary transformation that maps the computational basis onto the basis of fixed…

Quantum Physics · Physics 2025-11-25 María Laura Olivera-Atencio , Jesús Casado-Pascual , Denis Lacroix

Convergence of Q-learning has been the subject of extensive study for decades. Among the available techniques, the ordinary differential equation (ODE) method is particularly appealing as a general-purpose, off-the-shelf tool for…

Machine Learning · Computer Science 2026-05-12 Donghwan Lee , Hyunjun Na

The Koopman operator has emerged as a powerful tool for the analysis of nonlinear dynamical systems as it provides coordinate transformations to globally linearize the dynamics. While recent deep learning approaches have been useful in…

Dynamical Systems · Mathematics 2020-06-23 Shaowu Pan , Karthik Duraisamy

The use of target networks in deep reinforcement learning is a widely popular solution to mitigate the brittleness of semi-gradient approaches and stabilize learning. However, target networks notoriously require additional memory and delay…

Machine Learning · Computer Science 2026-03-02 Théo Vincent , Yogesh Tripathi , Tim Faust , Abdullah Akgül , Yaniv Oren , Melih Kandemir , Jan Peters , Carlo D'Eramo

In this paper, we propose Q-learning algorithms for continuous-time deterministic optimal control problems with Lipschitz continuous controls. Our method is based on a new class of Hamilton-Jacobi-Bellman (HJB) equations derived from…

Machine Learning · Computer Science 2020-10-28 Jeongho Kim , Jaeuk Shin , Insoon Yang

Value-based reinforcement learning (RL) methods like Q-learning have shown success in a variety of domains. One challenge in applying Q-learning to continuous-action RL problems, however, is the continuous action maximization (max-Q)…

Machine Learning · Computer Science 2020-03-03 Moonkyung Ryu , Yinlam Chow , Ross Anderson , Christian Tjandraatmadja , Craig Boutilier

We present a probabilistic viewpoint to multiple kernel learning unifying well-known regularised risk approaches and recent advances in approximate Bayesian inference relaxations. The framework proposes a general objective function suitable…

Machine Learning · Statistics 2012-06-08 Hannes Nickisch , Matthias Seeger

Evaluation of the Bellman functions is a difficult task. The exact Bellman functions of the dyadic Carleson Embedding Theorem 1.1 and the dyadic maximal operators are obtained in [3] and [4]. Actually, the same Bellman functions also work…

Classical Analysis and ODEs · Mathematics 2015-02-12 Jingguo Lai

Watkins' and Dayan's Q-learning is a model-free reinforcement learning algorithm that iteratively refines an estimate for the optimal action-value function of an MDP by stochastically "visiting" many state-ation pairs [Watkins and Dayan,…

Machine Learning · Computer Science 2021-08-09 Matthew T. Regehr , Alex Ayoub

The Koopman operator has become an essential tool for data-driven approximation of dynamical (control) systems, e.g., via extended dynamic mode decomposition. Despite its popularity, convergence results and, in particular, error bounds are…

Optimization and Control · Mathematics 2022-02-16 Feliks Nüske , Sebastian Peitz , Friedrich Philipp , Manuel Schaller , Karl Worthmann