English
Related papers

Related papers: Multi-Bellman operator for convergence of $Q$-lear…

200 papers

We introduce an abstract algorithm that aims to find the Bregman projection onto a closed convex set. As an application, the asymptotic behaviour of an iterative method for finding a fixed point of a quasi Bregman nonexpansive mapping with…

Functional Analysis · Mathematics 2013-09-26 Heinz H. Bauschke , Jiawei Chen , Xianfu Wang

$Q$-learning is one of the most fundamental reinforcement learning (RL) algorithms. Despite its widespread success in various applications, it is prone to overestimation bias in the $Q$-learning update. To address this issue, double…

Machine Learning · Computer Science 2026-01-13 Hyunjun Na , Donghwan Lee

Operator splitting techniques have recently gained popularity in convex optimization problems arising in various control fields. Being fixed-point iterations of nonexpansive operators, such methods suffer many well known downsides, which…

Optimization and Control · Mathematics 2020-04-01 Andreas Themelis , Panagiotis Patrinos

This paper addresses a learning problem for nonlinear dynamical systems with incorporating any specified dissipativity property. The nonlinear systems are described by the Koopman operator, which is a linear operator defined on the…

Systems and Control · Electrical Eng. & Systems 2019-11-12 Keita Hara , Masaki Inoue , Noboru Sebe

In reinforcement learning, the performance of learning agents is highly sensitive to the choice of time discretization. Agents acting at high frequencies have the best control opportunities, along with some drawbacks, such as possible…

Machine Learning · Computer Science 2022-11-22 Luca Sabbioni , Luca Al Daire , Lorenzo Bisi , Alberto Maria Metelli , Marcello Restelli

We propose a novel algorithmic framework for distributional reinforcement learning, based on learning finite-dimensional mean embeddings of return distributions. We derive several new algorithms for dynamic programming and…

This thesis explores a number of online machine learning algorithms. From a theoret- ical perspective, it assesses their employability for a particular function approximation problem where the analytical models fall short. Furthermore, it…

Machine Learning · Computer Science 2016-05-04 Ahmet Anil Pala

The fine-tuning of pre-trained large language models (LLMs) using reinforcement learning (RL) is generally formulated as direct policy optimization. This approach was naturally favored as it efficiently improves a pretrained LLM, seen as an…

We introduce a new algorithm for multi-objective reinforcement learning (MORL) with linear preferences, with the goal of enabling few-shot adaptation to new tasks. In MORL, the aim is to learn policies over multiple competing objectives…

Machine Learning · Computer Science 2019-11-07 Runzhe Yang , Xingyuan Sun , Karthik Narasimhan

Single-operator learning involves training a deep neural network to learn a specific operator, whereas recent work in multi-operator learning uses an operator embedding structure to train a single neural network on data from multiple…

Machine Learning · Computer Science 2025-06-16 Jingmin Sun , Zecheng Zhang , Hayden Schaeffer

Multi-Agent Reinforcement Learning involves agents that learn together in a shared environment, leading to emergent dynamics sensitive to initial conditions and parameter variations. A Dynamical Systems approach, which studies the evolution…

Multiagent Systems · Computer Science 2025-01-03 David Goll , Jobst Heitzig , Wolfram Barfuss

We study both the value function and Q-function formulation of the Linear Programming approach to Approximate Dynamic Programming. The approach is model-based and optimizes over a restricted function space to approximate the value function…

Systems and Control · Computer Science 2018-08-31 Paul N. Beuchat , Angelos Georghiou , John Lygeros

Q-learning (QL), a common reinforcement learning algorithm, suffers from over-estimation bias due to the maximization term in the optimal Bellman operator. This bias may lead to sub-optimal behavior. Double-Q-learning tackles this issue by…

Machine Learning · Computer Science 2021-04-21 Oren Peer , Chen Tessler , Nadav Merlis , Ron Meir

We propose a robust Q-learning algorithm for Markov decision processes under model uncertainty when each state-action pair is associated with a finite ambiguity set of candidate transition kernels. This finite-measure framework enables…

Optimization and Control · Mathematics 2026-02-10 Julian Sester , Cécile Decker

We attempt the use of a unitary operator to approximate the lattice Boltzmann collision operator. We use a modified amplitude encoding to bypass the renormalization that would have required classical processing at every step (thus eroding…

Quantum Physics · Physics 2026-01-08 Wael Itani , Katepalli R. Sreenivasan

This study investigates the application of quantum machine learning (QML) to approximate the nonlinear component of the collision operator within the quantum lattice Boltzmann method (QLBM). To achieve this, we train a variational quantum…

We study the convergence of a family of numerical integration methods where the numerical integral is formulated as a finite matrix approximation to a multiplication operator. For bounded functions, the convergence has already been…

Numerical Analysis · Mathematics 2023-03-28 Juha Sarmavuori , Simo Särkkä

This paper proposes a method to identify a Koopman model of a feedback-controlled system given a known controller. The Koopman operator allows a nonlinear system to be rewritten as an infinite-dimensional linear system by viewing it in…

Systems and Control · Electrical Eng. & Systems 2024-05-13 Steven Dahdah , James Richard Forbes

Discrete time stochastic optimal control problems and Markov decision processes (MDPs) are fundamental models for sequential decision-making under uncertainty and as such provide the mathematical framework underlying reinforcement learning…

Optimization and Control · Mathematics 2025-07-01 Arnulf Jentzen , Konrad Kleinberg , Thomas Kruse

Recently, reinforcement learning (RL) is receiving more and more attentions due to its successful demonstrations outperforming human performance in certain challenging tasks. In our recent paper `primal-dual Q-learning framework for LQR…

Optimization and Control · Mathematics 2018-11-22 Donghwan Lee , Jianghai Hu
‹ Prev 1 4 5 6 7 8 10 Next ›