English
Related papers

Related papers: A Small Gain Analysis of Single Timescale Actor Cr…

200 papers

We introduce a class of variational actor-critic algorithms based on a variational formulation over both the value function and the policy. The objective function of the variational formulation consists of two parts: one for maximizing the…

Machine Learning · Computer Science 2023-01-18 Yuhua Zhu , Lexing Ying

The hierarchical interaction between the actor and critic in actor-critic based reinforcement learning algorithms naturally lends itself to a game-theoretic interpretation. We adopt this viewpoint and model the actor and critic interaction…

Machine Learning · Computer Science 2021-09-28 Liyuan Zheng , Tanner Fiez , Zane Alumbaugh , Benjamin Chasnov , Lillian J. Ratliff

Recent multi-agent actor-critic methods have utilized centralized training with decentralized execution to address the non-stationarity of co-adapting agents. This training paradigm constrains learning to the centralized phase such that…

Multiagent Systems · Computer Science 2019-10-09 Kevin Corder , Manuel M. Vindiola , Keith Decker

Reinforcement learning algorithms are typically geared towards optimizing the expected return of an agent. However, in many practical applications, low variance in the return is desired to ensure the reliability of an algorithm. In this…

Machine Learning · Computer Science 2021-02-04 Arushi Jain , Gandharv Patil , Ayush Jain , Khimya Khetarpal , Doina Precup

We study the convergence of the actor-critic algorithm with nonlinear function approximation under a nonconvex-nonconcave primal-dual formulation. Stochastic gradient descent ascent is applied with an adaptive proximal term for robust…

Machine Learning · Computer Science 2022-03-01 Jing Dong , Li Shen , Yinggan Xu , Baoxiang Wang

A new Small-Gain Theorem is presented for general nonlinear control systems. The novelty of this research work is that vector Lyapunov functions and functionals are utilized to derive various input-to-output stability and input-to-state…

Optimization and Control · Mathematics 2009-04-07 Iasson Karafyllis , Zhong-Ping Jiang

Consider a decision maker who is responsible to collect observations so as to enhance his information in a speedy manner about an underlying phenomena of interest. The policies under which the decision maker selects sensing actions can be…

Information Theory · Computer Science 2015-06-12 Mohammad Naghshvar , Tara Javidi

Testing for change points in sequences of covariance matrices is an important and equally challenging problem in statistical methodology with applications in various fields. Motivated by the observation that even in cases where the ratio…

Statistics Theory · Mathematics 2026-01-14 Nina Dörnemann , Holger Dette

Vector autoregressive (VAR) models are widely used in multivariate time series analysis for describing the short-time dynamics of the data. The reduced-rank VAR models are of particular interest when dealing with high-dimensional and highly…

Statistics Theory · Mathematics 2023-05-02 Farida Enikeeva , Olga Klopp , Mathilde Rousselot

Stable distribution is one of the attractive models that well describes fat-tail behaviors and scaling phenomena in various scientific fields. The approach based upon the method of moments yields a simple procedure for estimating stable law…

Methodology · Statistics 2021-06-24 Shinji Kakinaka , Ken Umeno

Off-policy actor-critic algorithms have shown strong potential in deep reinforcement learning for continuous control tasks. Their success primarily comes from leveraging pessimistic state-action value function updates, which reduce function…

Machine Learning · Computer Science 2025-08-21 Bahareh Tasdighi , Nicklas Werge , Yi-Shan Wu , Melih Kandemir

Using the recent incremental modelling, it is shown that the trajectory of a sample in the phase space of soil mechanics in the vicinity of the critical state is not governed by the rigidity matrix, but by its variations. The…

Soft Condensed Matter · Physics 2007-05-23 P. Evesque

In a partially observed quantum or classical system the information that we cannot access results in our description of the system becoming mixed even if we have perfect initial knowledge. That is, if the system is quantum the conditional…

Quantum Physics · Physics 2009-11-11 Jay Gambetta , H. M. Wiseman

We present a computational strategy for reducing the sign problem in the evaluation of high dimensional integrals with non-positive definite weights. The method involves stochastic sampling with a positive semidefinite weight that is…

Computational Physics · Physics 2009-11-10 A G Moreira , S A Baeurle , G H Fredrickson

This paper is a continuation of the paper \cite{JL}, which focuses on exploring the global stability of nonlinear stochastic feedback systems on the nonnegative orthant driven by multiplicative white noise and presenting a couple of…

Dynamical Systems · Mathematics 2016-12-05 Jifa Jiang , Xiang Lv

This technical report is devoted to explaining how the actor loss of soft actor critic is obtained, as well as the associated gradient estimate. It gives the necessary mathematical background to derive all the presented equations, from the…

Machine Learning · Computer Science 2022-01-03 Thibault Lahire

Despite the popularity of the actor-critic method and the practical needs of collaborative policy training, existing works typically either overlook environmental heterogeneity or give up personalization altogether by training a single…

Machine Learning · Computer Science 2026-05-15 Leo Muxing Wang , Pengkun Yang , Lili Su

Small-gain conditions used in analysis of feedback interconnections are contraction conditions which imply certain stability properties. Such conditions are applied to a finite or infinite interval. In this paper we consider the case, when…

Dynamical Systems · Mathematics 2016-10-10 Petro Feketa , Humberto Stein Shiromoto , Sergey Dashkovskiy

Repeated small dynamic networks are integral to studies in evolutionary game theory, where networked public goods games offer novel insights into human behaviors. Building on these findings, it is necessary to develop a statistical model…

Applications · Statistics 2025-11-26 Hiroyasu Ando , Akihiro Nishi , Mark S. Handcock

Traditional video action detectors typically adopt the two-stage pipeline, where a person detector is first employed to generate actor boxes and then 3D RoIAlign is used to extract actor-specific features for classification. This detection…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Tao Wu , Mengqi Cao , Ziteng Gao , Gangshan Wu , Limin Wang