中文
相关论文

相关论文: A Small Gain Analysis of Single Timescale Actor Cr…

200 篇论文

Soft Actor-Critic (SAC) is widely used in practical applications and is now one of the most relevant off-policy online model-free reinforcement learning (RL) methods. The technique of n-step returns is known to increase the convergence…

机器学习 · 计算机科学 2025-12-16 Jakub Łyskawa , Jakub Lewandowski , Paweł Wawrzyński

In this paper, an alternative approximation to the innovation method is introduced for the parameter estimation of diffusion processes from partial and noisy observations. This is based on a convergent approximation to the first two…

最优化与控制 · 数学 2013-12-19 J. C. Jimenez

A sufficient condition for the stability of a system resulting from the interconnection of dynamical systems is given by the small gain theorem. Roughly speaking, to apply this theorem, it is required that the gains composition is…

动力系统 · 数学 2015-08-12 Humberto Stein Shiromoto , Vincent Andrieu , Christophe Prieur

When a learning algorithm reshapes the data distribution it trains on, the long-run behavior depends on the joint evolution of the policy, the value estimate, and the data distribution. We study finite-state actor-critic mean dynamics on…

动力系统 · 数学 2026-04-16 Vladyslav Prytula

Stable subordinators, and more general subordinators possessing power law probability tails, have been widely used in the context of subdiffusions, where particles get trapped or immobile in a number of time periods, called constant…

统计理论 · 数学 2020-05-11 Phillip Kerger , Kei Kobayashi

We introduce a reinforcement learning method for a class of non-Markov systems; our approach extends the actor-critic framework given by Rose et al. [New J. Phys. 23 013013 (2021)] for obtaining scaled cumulant generating functions…

统计力学 · 物理学 2026-03-09 Venkata D. Pamulaparthy , Rosemary J. Harris

Off-policy stochastic actor-critic methods rely on approximating the stochastic policy gradient in order to derive an optimal policy. One may also derive the optimal policy by approximating the action-value gradient. The use of action-value…

机器学习 · 统计学 2017-03-14 Yemi Okesanjo , Victor Kofia

We present a unified dynamical mean-field theory for stochastic self-organized critical models. We use a single site approximation and we include the details of different models by using effective parameters and constraints. We identify the…

统计力学 · 物理学 2009-10-28 Alessandro Vespignani , Stefano Zapperi

\Ac{MPC} and \ac{RL} are two powerful control strategies with, arguably, complementary advantages. In this work, we show how actor-critic \ac{RL} techniques can be leveraged to improve the performance of \ac{MPC}. The \ac{RL} critic is used…

系统与控制 · 电气工程与系统科学 2024-06-07 Rudolf Reiter , Andrea Ghezzi , Katrin Baumgärtner , Jasper Hoffmann , Robert D. McAllister , Moritz Diehl

Detecting changes in high-dimensional time series is difficult because it involves the comparison of probability densities that need to be estimated from finite samples. In this paper, we present the first feature extraction method tailored…

机器学习 · 计算机科学 2015-03-19 Duncan Blythe , Paul von Bünau , Frank Meinecke , Klaus-Robert Müller

For those seeking healthcare advice online, AI based dialogue agents capable of interacting with patients to perform automatic disease diagnosis are a viable option. This application necessitates efficient inquiry of relevant disease…

机器学习 · 计算机科学 2022-06-09 Weijie He , Ting Chen

We analyze a simple model of adaptive competition which captures essential features of a variety of adaptive competitive systems in the social and biological sciences. Each of N agents, at each time step of a game, joins one of two groups.…

adap-org · 物理学 2007-05-23 Robert Savit , Radu Manuca , Rick Riolo

A method of moment inequalities is used to derive the principle of minimum growth rate in multiplicatively interacting stochastic processes(MISPs). When a value of a power-law exponent at the tail of probability distribution function exists…

统计力学 · 物理学 2007-05-23 Akihiro Fujihara , Toshiya Ohtsuki , Hiroshi Yamamoto

The estimation of normalizing constants is a fundamental step in probabilistic model comparison. Sequential Monte Carlo methods may be used for this task and have the advantage of being inherently parallelizable. However, the standard…

机器学习 · 统计学 2016-08-16 Marco Fraccaro , Ulrich Paquet , Ole Winther

The stochastic actor oriented model (SAOM) is a method for modelling social interactions and social behaviour over time. It can be used to model drivers of dynamic interactions using both exogenous covariates and endogenous network…

统计方法学 · 统计学 2024-02-02 Giacomo Ceoldo , Tom A. B. Snijders , Ernst C. Wit

The usual passivity theorem considers a closed-loop, the direct chain of which consists of a strictly passive stable operator $H_{1}$, and the feedback chain of which consists of a passive operator $H_{2}$. Then the closed-loop is stable.…

最优化与控制 · 数学 2019-02-11 Henri Bourlès

Reinforcement learning algorithms are highly sensitive to the choice of hyperparameters, typically requiring significant manual effort to identify hyperparameters that perform well on a new domain. In this paper, we take a step towards…

Analysing stationary point databases to extract phenomenological rate constants can become time-consuming for systems with large potential energy barriers. In the present contribution we analyse several different approaches to this problem.…

软凝聚态物质 · 物理学 2009-11-11 Semen A. Trygubenko , David J. Wales

High-dimensional time series are a core ingredient of the statistical modeling toolkit, for which numerous estimation methods are known.But when observations are scarce or corrupted, the learning task becomes much harder.The question is:…

信号处理 · 电气工程与系统科学 2022-05-06 Guillaume Dalle , Yohann de Castro

Stationary stochastic processes with independent increments, of which the Poisson process is a prominent example, are widely used to describe real world events. With the basic assumption that a counting process is stationary and has…

概率论 · 数学 2018-11-20 Enzhi Li