中文
相关论文

相关论文: A Small Gain Analysis of Single Timescale Actor Cr…

200 篇论文

We introduce a class of variational actor-critic algorithms based on a variational formulation over both the value function and the policy. The objective function of the variational formulation consists of two parts: one for maximizing the…

机器学习 · 计算机科学 2023-01-18 Yuhua Zhu , Lexing Ying

The hierarchical interaction between the actor and critic in actor-critic based reinforcement learning algorithms naturally lends itself to a game-theoretic interpretation. We adopt this viewpoint and model the actor and critic interaction…

机器学习 · 计算机科学 2021-09-28 Liyuan Zheng , Tanner Fiez , Zane Alumbaugh , Benjamin Chasnov , Lillian J. Ratliff

Recent multi-agent actor-critic methods have utilized centralized training with decentralized execution to address the non-stationarity of co-adapting agents. This training paradigm constrains learning to the centralized phase such that…

多智能体系统 · 计算机科学 2019-10-09 Kevin Corder , Manuel M. Vindiola , Keith Decker

Reinforcement learning algorithms are typically geared towards optimizing the expected return of an agent. However, in many practical applications, low variance in the return is desired to ensure the reliability of an algorithm. In this…

机器学习 · 计算机科学 2021-02-04 Arushi Jain , Gandharv Patil , Ayush Jain , Khimya Khetarpal , Doina Precup

We study the convergence of the actor-critic algorithm with nonlinear function approximation under a nonconvex-nonconcave primal-dual formulation. Stochastic gradient descent ascent is applied with an adaptive proximal term for robust…

机器学习 · 计算机科学 2022-03-01 Jing Dong , Li Shen , Yinggan Xu , Baoxiang Wang

A new Small-Gain Theorem is presented for general nonlinear control systems. The novelty of this research work is that vector Lyapunov functions and functionals are utilized to derive various input-to-output stability and input-to-state…

最优化与控制 · 数学 2009-04-07 Iasson Karafyllis , Zhong-Ping Jiang

Consider a decision maker who is responsible to collect observations so as to enhance his information in a speedy manner about an underlying phenomena of interest. The policies under which the decision maker selects sensing actions can be…

信息论 · 计算机科学 2015-06-12 Mohammad Naghshvar , Tara Javidi

Testing for change points in sequences of covariance matrices is an important and equally challenging problem in statistical methodology with applications in various fields. Motivated by the observation that even in cases where the ratio…

统计理论 · 数学 2026-01-14 Nina Dörnemann , Holger Dette

Vector autoregressive (VAR) models are widely used in multivariate time series analysis for describing the short-time dynamics of the data. The reduced-rank VAR models are of particular interest when dealing with high-dimensional and highly…

统计理论 · 数学 2023-05-02 Farida Enikeeva , Olga Klopp , Mathilde Rousselot

Stable distribution is one of the attractive models that well describes fat-tail behaviors and scaling phenomena in various scientific fields. The approach based upon the method of moments yields a simple procedure for estimating stable law…

统计方法学 · 统计学 2021-06-24 Shinji Kakinaka , Ken Umeno

Off-policy actor-critic algorithms have shown strong potential in deep reinforcement learning for continuous control tasks. Their success primarily comes from leveraging pessimistic state-action value function updates, which reduce function…

机器学习 · 计算机科学 2025-08-21 Bahareh Tasdighi , Nicklas Werge , Yi-Shan Wu , Melih Kandemir

Using the recent incremental modelling, it is shown that the trajectory of a sample in the phase space of soil mechanics in the vicinity of the critical state is not governed by the rigidity matrix, but by its variations. The…

软凝聚态物质 · 物理学 2007-05-23 P. Evesque

In a partially observed quantum or classical system the information that we cannot access results in our description of the system becoming mixed even if we have perfect initial knowledge. That is, if the system is quantum the conditional…

量子物理 · 物理学 2009-11-11 Jay Gambetta , H. M. Wiseman

We present a computational strategy for reducing the sign problem in the evaluation of high dimensional integrals with non-positive definite weights. The method involves stochastic sampling with a positive semidefinite weight that is…

计算物理 · 物理学 2009-11-10 A G Moreira , S A Baeurle , G H Fredrickson

This paper is a continuation of the paper \cite{JL}, which focuses on exploring the global stability of nonlinear stochastic feedback systems on the nonnegative orthant driven by multiplicative white noise and presenting a couple of…

动力系统 · 数学 2016-12-05 Jifa Jiang , Xiang Lv

This technical report is devoted to explaining how the actor loss of soft actor critic is obtained, as well as the associated gradient estimate. It gives the necessary mathematical background to derive all the presented equations, from the…

机器学习 · 计算机科学 2022-01-03 Thibault Lahire

Despite the popularity of the actor-critic method and the practical needs of collaborative policy training, existing works typically either overlook environmental heterogeneity or give up personalization altogether by training a single…

机器学习 · 计算机科学 2026-05-15 Leo Muxing Wang , Pengkun Yang , Lili Su

Small-gain conditions used in analysis of feedback interconnections are contraction conditions which imply certain stability properties. Such conditions are applied to a finite or infinite interval. In this paper we consider the case, when…

动力系统 · 数学 2016-10-10 Petro Feketa , Humberto Stein Shiromoto , Sergey Dashkovskiy

Repeated small dynamic networks are integral to studies in evolutionary game theory, where networked public goods games offer novel insights into human behaviors. Building on these findings, it is necessary to develop a statistical model…

应用统计 · 统计学 2025-11-26 Hiroyasu Ando , Akihiro Nishi , Mark S. Handcock

Traditional video action detectors typically adopt the two-stage pipeline, where a person detector is first employed to generate actor boxes and then 3D RoIAlign is used to extract actor-specific features for classification. This detection…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Tao Wu , Mengqi Cao , Ziteng Gao , Gangshan Wu , Limin Wang