中文
相关论文

相关论文: A Small Gain Analysis of Single Timescale Actor Cr…

200 篇论文

In this work we study the problem of step size selection for numerical schemes, which guarantees that the numerical solution presents the same qualitative behavior as the original system of ordinary differential equations, by means of tools…

数值分析 · 数学 2015-05-13 Iasson Karafyllis , Lars Grune

This paper introduces the Active-Importance-Sampling Actor-Critic (AISAC) algorithm, an extension of the Actor-Critic framework for reducing variance in policy gradient estimation. AISAC optimizes the behavior policy to minimize gradient…

机器学习 · 计算机科学 2026-05-11 Majid Molaei , Gabor Paczolay , Matteo Papini , Alberto Maria Metelli , Marcello Restelli

Several recent works have focused on carrying out non-asymptotic convergence analyses for AC algorithms. Recently, a two-timescale critic-actor algorithm has been presented for the discounted cost setting in the look-up table case where the…

机器学习 · 计算机科学 2025-09-01 Prashansa Panda , Shalabh Bhatnagar

This paper presents a unification and a generalization of the small-gain theory subsuming a wide range of existing small-gain theorems. In particular, we introduce small-gain conditions that are necessary and sufficient to ensure…

最优化与控制 · 数学 2017-08-22 Navid Noroozi , Roman Geiselhart , Lars Grüne , Björn S. Rüffer , Fabian R. Wirth

In a wide range of applications, the stochastic properties of the observed time series change over time. The changes often occur gradually rather than abruptly: the properties are (approximately) constant for some time and then slowly start…

统计方法学 · 统计学 2015-04-03 Michael Vogt , Holger Dette

This paper is concerned with the detection of multiple change-points in the joint distribution of independent categorical variables. The procedures introduced rely on model selection and are based on a penalized least-squares criterion.…

统计理论 · 数学 2008-01-08 Nathalie Akakpo

Actor-critic methods constitute a central paradigm in reinforcement learning (RL), coupling policy evaluation with policy improvement. While effective across many domains, these methods rely on separate actor and critic networks, which…

机器学习 · 计算机科学 2025-09-26 Donghyeon Ki , Hee-Jun Ahn , Kyungyoon Kim , Byung-Jun Lee

This paper proposes a reinforcement learning (RL)-based backstepping control strategy to achieve fixed time consensus in nonlinear multi-agent systems with strict feedback dynamics. Agents exchange only output information with their…

系统与控制 · 电气工程与系统科学 2025-07-23 Aria Delshad , Maryam Babazadeh

In this paper, we establish the global optimality and convergence rate of an off-policy actor critic algorithm in the tabular setting without using density ratio to correct the discrepancy between the state distribution of the behavior…

机器学习 · 计算机科学 2025-02-07 Shangtong Zhang , Remi Tachet , Romain Laroche

We consider the problem of adaptive observer design in the settings when the system is allowed to be nonlinear in the parameters, and furthermore they are to satisfy additional feasibility constraints. A solution to the problem is proposed…

最优化与控制 · 数学 2014-12-18 I. Yu. Tyukin , P. A. Rogachev , H. Nijmeijer

Stein operators allow to characterise probability distributions via differential operators. Based on these characterisations, we develop a new method of point estimation for marginal parameters of strictly stationary and ergodic processes,…

统计理论 · 数学 2024-12-05 Bruno Ebner , Adrian Fischer , Robert E. Gaunt , Babette Picker , Yvik Swan

Instance segmentation is an important computer vision problem which remains challenging despite impressive recent advances due to deep learning-based methods. Given sufficient training data, fully supervised methods can yield excellent…

计算机视觉与模式识别 · 计算机科学 2023-03-22 Paul Hilt , Maedeh Zarvandi , Edgar Kaziakhmedov , Sourabh Bhide , Maria Leptin , Constantin Pape , Anna Kreshuk

We present a training framework for neural abstractive summarization based on actor-critic approaches from reinforcement learning. In the traditional neural network based methods, the objective is only to maximize the likelihood of the…

计算与语言 · 计算机科学 2018-08-16 Piji Li , Lidong Bing , Wai Lam

Policy gradient methods are reinforcement learning algorithms that adapt a parameterized policy by following a performance gradient estimate. Conventional policy gradient methods use Monte-Carlo techniques to estimate the gradient, which…

机器学习 · 计算机科学 2026-05-01 Mohammad Ghavamzadeh , Yaakov Engel , Michal Valko

We consider the superposition of a symmetric simple exclusion dynamics, speeded-up in time, with a spin-flip dynamics in a one-dimensional interval with periodic boundary conditions. We prove the large deviations principle for the empirical…

概率论 · 数学 2018-05-01 Jonathan Farfan , Claudio Landim , Kenkichi Tsunoda

We consider a model for a social network with N interacting social actors. This model is a system of interacting marked point processes in which each point process indicates the successive times in which a social actor expresses a…

概率论 · 数学 2024-05-29 Eva Löcherbach , Kádmo Laxa

We propose a penalized pseudo-likelihood criterion to estimate the graph of conditional dependencies in a discrete Markov random field that can be partially observed. We prove the convergence of the estimator in the case of a finite or…

统计方法学 · 统计学 2022-09-05 Florencia Leonardi , Rodrigo R. S. Carvalho

This paper studies a method, which has been proposed in the Physics literature by [8, 7, 10], for estimating the quasi-stationary distribution. In contrast to existing methods in eigenvector estimation, the method eliminates the need for…

概率论 · 数学 2014-01-03 Jose Blanchet , Peter Glynn , Shuheng Zheng

We present an approach to training neural networks to generate sequences using actor-critic methods from reinforcement learning (RL). Current log-likelihood training methods are limited by the discrepancy between their training and testing…

Actor-critic algorithms address the dual goals of reinforcement learning (RL), policy evaluation and improvement via two separate function approximators. The practicality of this approach comes at the expense of training instability, caused…

机器学习 · 计算机科学 2024-06-11 Bahareh Tasdighi , Abdullah Akgül , Manuel Haussmann , Kenny Kazimirzak Brink , Melih Kandemir