English
Related papers

Related papers: A Small Gain Analysis of Single Timescale Actor Cr…

200 papers

Current model-based reinforcement learning approaches use the model simply as a learned black-box simulator to augment the data for policy optimization or value function learning. In this paper, we show how to make more effective use of the…

Machine Learning · Computer Science 2020-05-19 Ignasi Clavera , Violet Fu , Pieter Abbeel

A classical problem for Markov chains is determining their stationary (or steady-state) distribution. This problem has an equally classical solution based on eigenvectors and linear equation systems. However, this approach does not scale to…

Systems and Control · Electrical Eng. & Systems 2023-01-20 Tobias Meggendorfer

We study the state consensus problem for linear shift-invariant discrete-time homogeneous multi-agent systems (MASs) over time-varying graphs. A novel approach based on the small gain theorem is proposed to design the consensus control…

Optimization and Control · Mathematics 2017-12-01 Ji-Lie Zhang , Xiang Chen , Guoxiang Gu

Weak consistency and asymptotic normality of the ordinary least-squares estimator in a linear regression with adaptive learning is derived when the crucial, so-called, `gain' parameter is estimated in a first step by nonlinear least squares…

Econometrics · Economics 2023-01-11 Alexander Mayer

The majority game, modelling a system of heterogeneous agents trying to behave in a similar way, is introduced and studied using methods of statistical mechanics. The stationary states of the game are given by the (local) minima of a…

Statistical Mechanics · Physics 2026-04-08 P. Kozlowski , M. Marsili

We prove a strong approximation result for the empirical process associated to a stationary sequence of real-valued random variables, under dependence conditions involving only indicators of half lines. This strong approximation result also…

Probability · Mathematics 2013-10-22 Jérôme Dedecker , Florence Merlevède , Emmanuel Rio

Optimization of parameterized policies for reinforcement learning (RL) is an important and challenging problem in artificial intelligence. Among the most common approaches are algorithms based on gradient ascent of a score function…

Machine Learning · Computer Science 2020-06-15 Sriram Srinivasan , Marc Lanctot , Vinicius Zambaldi , Julien Perolat , Karl Tuyls , Remi Munos , Michael Bowling

Existing work on risk-sensitive reinforcement learning - both for symmetric and downside risk measures - has typically used direct Monte-Carlo estimation of policy gradients. While this approach yields unbiased gradient estimates, it also…

Machine Learning · Computer Science 2020-07-09 Thomas Spooner , Rahul Savani

Policy gradient algorithms typically combine discounted future rewards with an estimated value function, to compute the direction and magnitude of parameter updates. However, for most Reinforcement Learning tasks, humans can provide…

Machine Learning · Computer Science 2019-04-09 Ishan Durugkar , Matthew Hausknecht , Adith Swaminathan , Patrick MacAlpine

While there has been substantial success for solving continuous control with actor-critic methods, simpler critic-only methods such as Q-learning find limited application in the associated high-dimensional action spaces. However, most…

Machine Learning · Computer Science 2023-09-27 Tim Seyde , Peter Werner , Wilko Schwarting , Igor Gilitschenski , Martin Riedmiller , Daniela Rus , Markus Wulfmeier

This paper proposes the Cooperative Soft Actor Critic (CSAC) method of enabling consecutive reinforcement learning agents to cooperatively solve a long time horizon multi-stage task. This method is achieved by modifying the policy of each…

Machine Learning · Computer Science 2020-07-02 Jordan Erskine , Chris Lehnert

We propose an informal test for stationarity in a time series which checks for the compatibility of nonlinear approximations to the dynamics made in different segments of the sequence. The segments are compared directly, rather than via…

chao-dyn · Physics 2009-10-31 Thomas Schreiber

We prove under commonly used assumptions the convergence of actor-critic reinforcement learning algorithms, which simultaneously learn a policy function, the actor, and a value function, the critic. Both functions can be deep neural…

Machine Learning · Computer Science 2020-12-03 Markus Holzleitner , Lukas Gruber , José Arjona-Medina , Johannes Brandstetter , Sepp Hochreiter

In this paper, we investigate the issue of error accumulation in critic networks updated via pessimistic temporal difference objectives. We show that the critic approximation error can be approximated via a recursive fixed-point model…

Machine Learning · Computer Science 2024-03-05 Michal Nauman , Mateusz Ostaszewski , Marek Cygan

The present paper proposes a new treatment effects estimator that is valid when the number of time periods is small, and the parallel trends condition holds conditional on covariates and unobserved heterogeneity in the form of interactive…

Econometrics · Economics 2023-06-16 Nicholas Brown , Kyle Butts , Joakim Westerlund

A networked output feedback loop subject to packetized transmissions of the output signal is considered. Based on the small gain theorem, an easy-to-use stability criterion covering two important cases is presented. In the first case a…

Systems and Control · Electrical Eng. & Systems 2021-08-24 Martin Steinberger , Martin Horn

In this paper, we propose a distributed off-policy actor critic method to solve multi-agent reinforcement learning problems. Specifically, we assume that all agents keep local estimates of the global optimal policy parameter and update…

Machine Learning · Computer Science 2019-03-25 Yan Zhang , Michael M. Zavlanos

We review some techniques from non-linear analysis in order to investigate critical paths for the action functional in the calculus of variations applied to physics. Previous attempts to analyse when these are minima ex- ist, but mainly…

Mathematical Physics · Physics 2013-03-22 E. López , A. Molgado , J. A. Vallejo

By using an parametric value function to replace the Monte-Carlo rollouts for value estimation, the actor-critic (AC) algorithms can reduce the variance of stochastic policy gradient so that to improve the convergence rate. While existing…

Machine Learning · Computer Science 2024-08-19 Yanjie Dong , Haijun Zhang , Gang Wang , Shisheng Cui , Xiping Hu

We propose a novel independent and payoff-based learning framework for stochastic games that is model-free, game-agnostic, and gradient-free. The learning dynamics follow a best-response-type actor-critic architecture, where agents update…

Machine Learning · Computer Science 2026-02-03 Ahmed Said Donmez , Yuksel Arslantas , Muhammed O. Sayin
‹ Prev 1 8 9 10 Next ›