中文
相关论文

相关论文: A Small Gain Analysis of Single Timescale Actor Cr…

200 篇论文

In this paper, we propose a second-order deterministic actor-critic framework in reinforcement learning that extends the classical deterministic policy gradient method to exploit curvature information of the performance function. Building…

机器学习 · 计算机科学 2025-11-13 Arash Bahari Kordabad , Dean Brandner , Sebastien Gros , Sergio Lucia , Sadegh Soudjani

Efficient utilization of the replay buffer plays a significant role in the off-policy actor-critic reinforcement learning (RL) algorithms used for model-free control policy synthesis for complex dynamical systems. We propose a method for…

机器学习 · 计算机科学 2024-02-13 Nikhil Kumar Singh , Indranil Saha

Soft Actor-Critic is a state-of-the-art reinforcement learning algorithm for continuous action settings that is not applicable to discrete action settings. Many important settings involve discrete actions, however, and so here we derive an…

机器学习 · 计算机科学 2019-10-21 Petros Christodoulou

We introduce a simple extension of the minority game in which the market rewards contrarian (resp. trend-following) strategies when it is far from (resp. close to) efficiency. The model displays a smooth crossover from a regime where…

无序系统与神经网络 · 物理学 2009-11-10 A. De Martino , I. Giardina , M. Marsili , A. Tedeschi

We consider the problem of learning the parameters of a $N$-dimensional stochastic linear dynamics under both full and partial observations from a single trajectory of time $T$. We introduce and analyze a new estimator that achieves a small…

机器学习 · 统计学 2025-12-08 Minh Vu , Andrey Y. Lokhov , Marc Vuffray

We propose a novel actor-critic algorithm with guaranteed convergence to an optimal policy for a discounted reward Markov decision process. The actor incorporates a descent direction that is motivated by the solution of a certain non-linear…

机器学习 · 计算机科学 2015-07-30 Prashanth L. A. , H. L. Prasad , Shalabh Bhatnagar , Prakash Chandra

A small-gain approach is proposed to analyze closed-loop stability of linear diffusion-reaction systems under finite-dimensional observer-based state feedback control. For this, the decomposition of the infinite-dimensional system into a…

系统与控制 · 电气工程与系统科学 2022-02-14 Lars Grüne , Thomas Meurer

As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical question emerges: How does the interaction between researchers and agents affect the results?…

We propose a new algorithm, Mean Actor-Critic (MAC), for discrete-action continuous-state reinforcement learning. MAC is a policy gradient algorithm that uses the agent's explicit representation of all action values to estimate the gradient…

Optimal control problems with free terminal time present many challenges including nonsmooth and discontinuous control laws, irregular value functions, many local optima, and the curse of dimensionality. To overcome these issues, we propose…

最优化与控制 · 数学 2022-08-08 Evan Burton , Tenavi Nakamura-Zimmerer , Qi Gong , Wei Kang

We estimate the time a point or set, respectively, requires to approach the attractor of a radially symmetric gradient type stochastic differential equation driven by small noise. Here, both of these times tend to infinity as the noise gets…

概率论 · 数学 2018-06-07 Isabell Vorkastner

Subject to reasonable conditions, in large population stochastic dynamics games, where the agents are coupled by the system's mean field (i.e. the state distribution of the generic agent) through their nonlinear dynamics and their nonlinear…

最优化与控制 · 数学 2019-05-28 Nevroz Sen , Peter E. Caines

Under a complex technical condition, similar to such used in extreme value theory, we find the rate q(\epsilon)^{-1} at which a stochastic process with stationary increments \xi should be sampled, for the sampled process \xi(\lfloor\cdot…

概率论 · 数学 2007-05-23 J. M. P. Albin

This paper develops a unified and computationally efficient method for change-point estimation along the time dimension in a non-stationary spatio-temporal process. By modeling a non-stationary spatio-temporal process as a piecewise…

统计方法学 · 统计学 2023-10-09 Zifeng Zhao , Ting Fung Ma , Wai Leong Ng , Chun Yip Yau

The cyclic feedback interconnection of $n$ subsystems is the basic building block of control theory. Many robust stability tools have been developed for this interconnection. Two notable examples are the small gain theorem and the Secant…

最优化与控制 · 数学 2023-05-04 Richard Pates

A parameter estimation method is devised for a slow-fast stochastic dynamical system, where often only the slow component is observable. By using the observations only on the slow component, the system parameters are estimated by working on…

动力系统 · 数学 2013-03-20 Jian Ren , Jinqiao Duan

Off-Policy Actor-Critic (Off-PAC) methods have proven successful in a variety of continuous control tasks. Normally, the critic's action-value function is updated using temporal-difference, and the critic in turn provides a loss for the…

机器学习 · 计算机科学 2020-11-03 Wei Zhou , Yiying Li , Yongxin Yang , Huaimin Wang , Timothy M. Hospedales

Many policy gradient methods are variants of Actor-Critic (AC), where a value function (critic) is learned to facilitate updating the parameterized policy (actor). The update to the actor involves a log-likelihood update weighted by the…

机器学习 · 计算机科学 2023-03-02 Samuel Neumann , Sungsu Lim , Ajin Joseph , Yangchen Pan , Adam White , Martha White

In a wide range of applications, the stochastic properties of the observed time series change over time. The changes often occur gradually rather than abruptly: the prop- erties are (approximately) constant for some time and then slowly…

统计方法学 · 统计学 2014-03-18 Michael Vogt , Holger Dette

This paper introduces small-gain sufficient conditions for $2$-contraction of feedback interconnected systems, on the basis of individual gains of suitable subsystems arising from a modular decomposition of the second additive compound…

系统与控制 · 电气工程与系统科学 2023-07-03 David Angeli , Davide Martini , Giacomo Innocenti , Alberto Tesi