Related papers: A Small Gain Analysis of Single Timescale Actor Cr…
In this paper, we propose a second-order deterministic actor-critic framework in reinforcement learning that extends the classical deterministic policy gradient method to exploit curvature information of the performance function. Building…
Efficient utilization of the replay buffer plays a significant role in the off-policy actor-critic reinforcement learning (RL) algorithms used for model-free control policy synthesis for complex dynamical systems. We propose a method for…
Soft Actor-Critic is a state-of-the-art reinforcement learning algorithm for continuous action settings that is not applicable to discrete action settings. Many important settings involve discrete actions, however, and so here we derive an…
We introduce a simple extension of the minority game in which the market rewards contrarian (resp. trend-following) strategies when it is far from (resp. close to) efficiency. The model displays a smooth crossover from a regime where…
We consider the problem of learning the parameters of a $N$-dimensional stochastic linear dynamics under both full and partial observations from a single trajectory of time $T$. We introduce and analyze a new estimator that achieves a small…
We propose a novel actor-critic algorithm with guaranteed convergence to an optimal policy for a discounted reward Markov decision process. The actor incorporates a descent direction that is motivated by the solution of a certain non-linear…
A small-gain approach is proposed to analyze closed-loop stability of linear diffusion-reaction systems under finite-dimensional observer-based state feedback control. For this, the decomposition of the infinite-dimensional system into a…
As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical question emerges: How does the interaction between researchers and agents affect the results?…
We propose a new algorithm, Mean Actor-Critic (MAC), for discrete-action continuous-state reinforcement learning. MAC is a policy gradient algorithm that uses the agent's explicit representation of all action values to estimate the gradient…
Optimal control problems with free terminal time present many challenges including nonsmooth and discontinuous control laws, irregular value functions, many local optima, and the curse of dimensionality. To overcome these issues, we propose…
We estimate the time a point or set, respectively, requires to approach the attractor of a radially symmetric gradient type stochastic differential equation driven by small noise. Here, both of these times tend to infinity as the noise gets…
Subject to reasonable conditions, in large population stochastic dynamics games, where the agents are coupled by the system's mean field (i.e. the state distribution of the generic agent) through their nonlinear dynamics and their nonlinear…
Under a complex technical condition, similar to such used in extreme value theory, we find the rate q(\epsilon)^{-1} at which a stochastic process with stationary increments \xi should be sampled, for the sampled process \xi(\lfloor\cdot…
This paper develops a unified and computationally efficient method for change-point estimation along the time dimension in a non-stationary spatio-temporal process. By modeling a non-stationary spatio-temporal process as a piecewise…
The cyclic feedback interconnection of $n$ subsystems is the basic building block of control theory. Many robust stability tools have been developed for this interconnection. Two notable examples are the small gain theorem and the Secant…
A parameter estimation method is devised for a slow-fast stochastic dynamical system, where often only the slow component is observable. By using the observations only on the slow component, the system parameters are estimated by working on…
Off-Policy Actor-Critic (Off-PAC) methods have proven successful in a variety of continuous control tasks. Normally, the critic's action-value function is updated using temporal-difference, and the critic in turn provides a loss for the…
Many policy gradient methods are variants of Actor-Critic (AC), where a value function (critic) is learned to facilitate updating the parameterized policy (actor). The update to the actor involves a log-likelihood update weighted by the…
In a wide range of applications, the stochastic properties of the observed time series change over time. The changes often occur gradually rather than abruptly: the prop- erties are (approximately) constant for some time and then slowly…
This paper introduces small-gain sufficient conditions for $2$-contraction of feedback interconnected systems, on the basis of individual gains of suitable subsystems arising from a modular decomposition of the second additive compound…