Related papers: A Small Gain Analysis of Single Timescale Actor Cr…
Soft Actor-Critic (SAC) is widely used in practical applications and is now one of the most relevant off-policy online model-free reinforcement learning (RL) methods. The technique of n-step returns is known to increase the convergence…
In this paper, an alternative approximation to the innovation method is introduced for the parameter estimation of diffusion processes from partial and noisy observations. This is based on a convergent approximation to the first two…
A sufficient condition for the stability of a system resulting from the interconnection of dynamical systems is given by the small gain theorem. Roughly speaking, to apply this theorem, it is required that the gains composition is…
When a learning algorithm reshapes the data distribution it trains on, the long-run behavior depends on the joint evolution of the policy, the value estimate, and the data distribution. We study finite-state actor-critic mean dynamics on…
Stable subordinators, and more general subordinators possessing power law probability tails, have been widely used in the context of subdiffusions, where particles get trapped or immobile in a number of time periods, called constant…
We introduce a reinforcement learning method for a class of non-Markov systems; our approach extends the actor-critic framework given by Rose et al. [New J. Phys. 23 013013 (2021)] for obtaining scaled cumulant generating functions…
Off-policy stochastic actor-critic methods rely on approximating the stochastic policy gradient in order to derive an optimal policy. One may also derive the optimal policy by approximating the action-value gradient. The use of action-value…
We present a unified dynamical mean-field theory for stochastic self-organized critical models. We use a single site approximation and we include the details of different models by using effective parameters and constraints. We identify the…
\Ac{MPC} and \ac{RL} are two powerful control strategies with, arguably, complementary advantages. In this work, we show how actor-critic \ac{RL} techniques can be leveraged to improve the performance of \ac{MPC}. The \ac{RL} critic is used…
Detecting changes in high-dimensional time series is difficult because it involves the comparison of probability densities that need to be estimated from finite samples. In this paper, we present the first feature extraction method tailored…
For those seeking healthcare advice online, AI based dialogue agents capable of interacting with patients to perform automatic disease diagnosis are a viable option. This application necessitates efficient inquiry of relevant disease…
We analyze a simple model of adaptive competition which captures essential features of a variety of adaptive competitive systems in the social and biological sciences. Each of N agents, at each time step of a game, joins one of two groups.…
A method of moment inequalities is used to derive the principle of minimum growth rate in multiplicatively interacting stochastic processes(MISPs). When a value of a power-law exponent at the tail of probability distribution function exists…
The estimation of normalizing constants is a fundamental step in probabilistic model comparison. Sequential Monte Carlo methods may be used for this task and have the advantage of being inherently parallelizable. However, the standard…
The stochastic actor oriented model (SAOM) is a method for modelling social interactions and social behaviour over time. It can be used to model drivers of dynamic interactions using both exogenous covariates and endogenous network…
The usual passivity theorem considers a closed-loop, the direct chain of which consists of a strictly passive stable operator $H_{1}$, and the feedback chain of which consists of a passive operator $H_{2}$. Then the closed-loop is stable.…
Reinforcement learning algorithms are highly sensitive to the choice of hyperparameters, typically requiring significant manual effort to identify hyperparameters that perform well on a new domain. In this paper, we take a step towards…
Analysing stationary point databases to extract phenomenological rate constants can become time-consuming for systems with large potential energy barriers. In the present contribution we analyse several different approaches to this problem.…
High-dimensional time series are a core ingredient of the statistical modeling toolkit, for which numerous estimation methods are known.But when observations are scarce or corrupted, the learning task becomes much harder.The question is:…
Stationary stochastic processes with independent increments, of which the Poisson process is a prominent example, are widely used to describe real world events. With the basic assumption that a counting process is stationary and has…