Related papers: A Small Gain Analysis of Single Timescale Actor Cr…
We introduce a class of variational actor-critic algorithms based on a variational formulation over both the value function and the policy. The objective function of the variational formulation consists of two parts: one for maximizing the…
The hierarchical interaction between the actor and critic in actor-critic based reinforcement learning algorithms naturally lends itself to a game-theoretic interpretation. We adopt this viewpoint and model the actor and critic interaction…
Recent multi-agent actor-critic methods have utilized centralized training with decentralized execution to address the non-stationarity of co-adapting agents. This training paradigm constrains learning to the centralized phase such that…
Reinforcement learning algorithms are typically geared towards optimizing the expected return of an agent. However, in many practical applications, low variance in the return is desired to ensure the reliability of an algorithm. In this…
We study the convergence of the actor-critic algorithm with nonlinear function approximation under a nonconvex-nonconcave primal-dual formulation. Stochastic gradient descent ascent is applied with an adaptive proximal term for robust…
A new Small-Gain Theorem is presented for general nonlinear control systems. The novelty of this research work is that vector Lyapunov functions and functionals are utilized to derive various input-to-output stability and input-to-state…
Consider a decision maker who is responsible to collect observations so as to enhance his information in a speedy manner about an underlying phenomena of interest. The policies under which the decision maker selects sensing actions can be…
Testing for change points in sequences of covariance matrices is an important and equally challenging problem in statistical methodology with applications in various fields. Motivated by the observation that even in cases where the ratio…
Vector autoregressive (VAR) models are widely used in multivariate time series analysis for describing the short-time dynamics of the data. The reduced-rank VAR models are of particular interest when dealing with high-dimensional and highly…
Stable distribution is one of the attractive models that well describes fat-tail behaviors and scaling phenomena in various scientific fields. The approach based upon the method of moments yields a simple procedure for estimating stable law…
Off-policy actor-critic algorithms have shown strong potential in deep reinforcement learning for continuous control tasks. Their success primarily comes from leveraging pessimistic state-action value function updates, which reduce function…
Using the recent incremental modelling, it is shown that the trajectory of a sample in the phase space of soil mechanics in the vicinity of the critical state is not governed by the rigidity matrix, but by its variations. The…
In a partially observed quantum or classical system the information that we cannot access results in our description of the system becoming mixed even if we have perfect initial knowledge. That is, if the system is quantum the conditional…
We present a computational strategy for reducing the sign problem in the evaluation of high dimensional integrals with non-positive definite weights. The method involves stochastic sampling with a positive semidefinite weight that is…
This paper is a continuation of the paper \cite{JL}, which focuses on exploring the global stability of nonlinear stochastic feedback systems on the nonnegative orthant driven by multiplicative white noise and presenting a couple of…
This technical report is devoted to explaining how the actor loss of soft actor critic is obtained, as well as the associated gradient estimate. It gives the necessary mathematical background to derive all the presented equations, from the…
Despite the popularity of the actor-critic method and the practical needs of collaborative policy training, existing works typically either overlook environmental heterogeneity or give up personalization altogether by training a single…
Small-gain conditions used in analysis of feedback interconnections are contraction conditions which imply certain stability properties. Such conditions are applied to a finite or infinite interval. In this paper we consider the case, when…
Repeated small dynamic networks are integral to studies in evolutionary game theory, where networked public goods games offer novel insights into human behaviors. Building on these findings, it is necessary to develop a statistical model…
Traditional video action detectors typically adopt the two-stage pipeline, where a person detector is first employed to generate actor boxes and then 3D RoIAlign is used to extract actor-specific features for classification. This detection…