相关论文: Global attractors and fast-slow reduction for fini…
Actor-critic style two-time-scale algorithms are one of the most popular methods in reinforcement learning, and have seen great empirical success. However, their performance is not completely understood theoretically. In this paper, we…
In this paper, we establish the global optimality and convergence rate of an off-policy actor critic algorithm in the tabular setting without using density ratio to correct the discrepancy between the state distribution of the behavior…
We study the global convergence and global optimality of actor-critic, one of the most popular families of reinforcement learning algorithms. While most existing works on actor-critic employ bi-level or two-timescale updates, we focus on…
Motivated by applications in risk-sensitive reinforcement learning, we study mean-variance optimization in a discounted reward Markov Decision Process (MDP). Specifically, we analyze a Temporal Difference (TD) learning algorithm with linear…
We consider the reinforcement learning problem for partially observed Markov decision processes (POMDPs) with large or even countably infinite state spaces, where the controller has access to only noisy observations of the underlying…
We deal with a class of parabolic nonlinear evolution equations with state-dependent delay. This class covers several important PDE models arising in biology. We first prove well-posedness in a certain space of functions which are Lipschitz…
In a recent article, we introduced the concept of streams and graphs of a semiflow. An important related concept is the one of semiflow with {\em compact dynamics}, which we defined as a semiflow $F$ with a {\em compact global trapping…
We address, in a three-dimensional spatial setting, both the viscous and the standard Cahn-Hilliard equation with a nonconstant mobility coefficient. As it was shown in J.W. Barrett and J.W. Blowey, Math. Comp., 68 (1999), 487-517, one…
Actor-critic algorithms are widely used in reinforcement learning, but are challenging to mathematically analyse due to the online arrival of non-i.i.d. data samples. The distribution of the data samples dynamically changes as the model is…
Reinforcement learning (RL) has gained attention for aligning large language models (LLMs) via reinforcement learning from human feedback (RLHF). The actor-only variants of Proximal Policy Optimization (PPO) are widely applied for their…
Global dynamics of the diffusive Hindmarsh-Rose equations with memristor as a new proposed model for neuron dynamics are investigated in this paper. We prove the existence and regularity of a global attractor for the solution semiflow…
Despite the popularity of the actor-critic method and the practical needs of collaborative policy training, existing works typically either overlook environmental heterogeneity or give up personalization altogether by training a single…
We consider piecewise linear discrete time macroeconomic models, which possess a continuum of equilibrium states. These systems are obtained by replacing rational inflation expectations with a boundedly rational, and genuinely sticky,…
The long-time behavior of the solutions for a non-isothermal model in superfluidity is investigated. The model describes the transition between the normal and the superfluid phase in liquid 4He by means of a non-linear differential system,…
Multi-agent reinforcement learning has been successfully applied to a number of challenging problems. Despite these empirical successes, theoretical understanding of different algorithms is lacking, primarily due to the curse of…
This paper aims at distributed multi-agent convex optimization where the communications network among the agents are presented by a random sequence of possibly state-dependent weighted graphs. This is the first work to consider both random…
We establish the well-posedness of a strongly damped semilinear wave equation equipped with nonlinear hyperbolic dynamic boundary conditions. Results are carried out with the presence of a parameter distinguishing whether the underlying…
In reinforcement learning for partially observable environments, many successful algorithms have been developed within the asymmetric learning paradigm. This paradigm leverages additional state information available at training time for…
We analyze the global convergence of the single-timescale actor-critic (AC) algorithm for the infinite-horizon discounted Markov Decision Processes (MDPs) with finite state spaces. To this end, we introduce an elegant analytical framework…
The wave equation with energy critical sources and nonlinear damping defined on a 3D bounded domain is considered. It is shown that the resulting dynamical system admits a global attractor. Under the additional assumption of strong…