中文
相关论文

相关论文: Global attractors and fast-slow reduction for fini…

200 篇论文

Actor-critic style two-time-scale algorithms are one of the most popular methods in reinforcement learning, and have seen great empirical success. However, their performance is not completely understood theoretically. In this paper, we…

机器学习 · 计算机科学 2022-02-22 Sajad Khodadadian , Thinh T. Doan , Justin Romberg , Siva Theja Maguluri

In this paper, we establish the global optimality and convergence rate of an off-policy actor critic algorithm in the tabular setting without using density ratio to correct the discrepancy between the state distribution of the behavior…

机器学习 · 计算机科学 2025-02-07 Shangtong Zhang , Remi Tachet , Romain Laroche

We study the global convergence and global optimality of actor-critic, one of the most popular families of reinforcement learning algorithms. While most existing works on actor-critic employ bi-level or two-timescale updates, we focus on…

机器学习 · 计算机科学 2021-06-15 Zuyue Fu , Zhuoran Yang , Zhaoran Wang

Motivated by applications in risk-sensitive reinforcement learning, we study mean-variance optimization in a discounted reward Markov Decision Process (MDP). Specifically, we analyze a Temporal Difference (TD) learning algorithm with linear…

机器学习 · 计算机科学 2025-03-13 Tejaram Sangadi , L. A. Prashanth , Krishna Jagannathan

We consider the reinforcement learning problem for partially observed Markov decision processes (POMDPs) with large or even countably infinite state spaces, where the controller has access to only noisy observations of the underlying…

机器学习 · 计算机科学 2023-07-20 Semih Cayci , Niao He , R. Srikant

We deal with a class of parabolic nonlinear evolution equations with state-dependent delay. This class covers several important PDE models arising in biology. We first prove well-posedness in a certain space of functions which are Lipschitz…

偏微分方程分析 · 数学 2016-03-22 Igor Chueshov , Alexander Rezounenko

In a recent article, we introduced the concept of streams and graphs of a semiflow. An important related concept is the one of semiflow with {\em compact dynamics}, which we defined as a semiflow $F$ with a {\em compact global trapping…

动力系统 · 数学 2025-03-05 Roberto De Leo , James A. Yorke

We address, in a three-dimensional spatial setting, both the viscous and the standard Cahn-Hilliard equation with a nonconstant mobility coefficient. As it was shown in J.W. Barrett and J.W. Blowey, Math. Comp., 68 (1999), 487-517, one…

偏微分方程分析 · 数学 2009-11-13 Giulio Schimperna

Actor-critic algorithms are widely used in reinforcement learning, but are challenging to mathematically analyse due to the online arrival of non-i.i.d. data samples. The distribution of the data samples dynamically changes as the model is…

机器学习 · 计算机科学 2023-09-20 Ziheng Wang , Justin Sirignano

Reinforcement learning (RL) has gained attention for aligning large language models (LLMs) via reinforcement learning from human feedback (RLHF). The actor-only variants of Proximal Policy Optimization (PPO) are widely applied for their…

最优化与控制 · 数学 2025-12-19 Yin Liu , Qiming Dai , Junyu Zhang , Zaiwen Wen

Global dynamics of the diffusive Hindmarsh-Rose equations with memristor as a new proposed model for neuron dynamics are investigated in this paper. We prove the existence and regularity of a global attractor for the solution semiflow…

偏微分方程分析 · 数学 2022-08-23 Yuncheng You

Despite the popularity of the actor-critic method and the practical needs of collaborative policy training, existing works typically either overlook environmental heterogeneity or give up personalization altogether by training a single…

机器学习 · 计算机科学 2026-05-15 Leo Muxing Wang , Pengkun Yang , Lili Su

We consider piecewise linear discrete time macroeconomic models, which possess a continuum of equilibrium states. These systems are obtained by replacing rational inflation expectations with a boundedly rational, and genuinely sticky,…

动力系统 · 数学 2017-11-22 Pavel Krejci , Harbir Lamba , Dmitrii Rachinskii

The long-time behavior of the solutions for a non-isothermal model in superfluidity is investigated. The model describes the transition between the normal and the superfluid phase in liquid 4He by means of a non-linear differential system,…

数学物理 · 物理学 2011-02-08 Alessia Berti , Valeria Berti , Ivana Bochicchio

Multi-agent reinforcement learning has been successfully applied to a number of challenging problems. Despite these empirical successes, theoretical understanding of different algorithms is lacking, primarily due to the curse of…

机器学习 · 计算机科学 2021-12-28 Yuwei Luo , Zhuoran Yang , Zhaoran Wang , Mladen Kolar

This paper aims at distributed multi-agent convex optimization where the communications network among the agents are presented by a random sequence of possibly state-dependent weighted graphs. This is the first work to consider both random…

系统与控制 · 电气工程与系统科学 2024-12-31 Seyyed Shaho Alaviani , Atul Kelkar

We establish the well-posedness of a strongly damped semilinear wave equation equipped with nonlinear hyperbolic dynamic boundary conditions. Results are carried out with the presence of a parameter distinguishing whether the underlying…

偏微分方程分析 · 数学 2016-03-23 P. Jameson Graber , Joseph L. Shomberg

In reinforcement learning for partially observable environments, many successful algorithms have been developed within the asymmetric learning paradigm. This paradigm leverages additional state information available at training time for…

机器学习 · 计算机科学 2025-09-09 Gaspard Lambrechts , Damien Ernst , Aditya Mahajan

We analyze the global convergence of the single-timescale actor-critic (AC) algorithm for the infinite-horizon discounted Markov Decision Processes (MDPs) with finite state spaces. To this end, we introduce an elegant analytical framework…

机器学习 · 计算机科学 2025-06-05 Navdeep Kumar , Priyank Agrawal , Giorgia Ramponi , Kfir Yehuda Levy , Shie Mannor

The wave equation with energy critical sources and nonlinear damping defined on a 3D bounded domain is considered. It is shown that the resulting dynamical system admits a global attractor. Under the additional assumption of strong…

动力系统 · 数学 2025-11-07 Irena Lasiecka , Vando Narciso
‹ 上一页 1 2 3 10 下一页 ›