中文
相关论文

相关论文: How are policy gradient methods affected by the li…

200 篇论文

Typical properties of computing circuits composed of noisy logical gates are studied using the statistical physics methodology. A growth model that gives rise to typical random Boolean functions is mapped onto a layered Ising spin system,…

无序系统与神经网络 · 物理学 2015-05-18 Alexander Mozeika , David Saad , Jack Raymond

We explore reinforcement learning methods for finding the optimal policy in the linear quadratic regulator (LQR) problem. In particular, we consider the convergence of policy gradient methods in the setting of known and unknown parameters.…

机器学习 · 计算机科学 2021-06-25 Ben Hambly , Renyuan Xu , Huining Yang

Accurate state estimation requires careful consideration of uncertainty surrounding the process and measurement models; these characteristics are usually not well-known and need an experienced designer to select the covariance matrices. An…

机器学习 · 统计学 2025-07-18 Pardha Sai Krishna Ala , Ameya Salvi , Venkat Krovi , Matthias Schmid

Stochastic gradient optimization is the dominant learning paradigm for a variety of scenarios, from classical supervised learning to modern self-supervised learning. We consider stochastic gradient algorithms for learning problems whose…

机器学习 · 统计学 2025-08-29 Facheng Yu , Ronak Mehta , Alex Luedtke , Zaid Harchaoui

Stochastic systems have a control-theoretic interpretation in which noise plays the role of control. In the weak-noise limit, relevant at low temperatures or in large populations, this leads to a precise mathematical mapping: the most…

分子网络 · 定量生物学 2025-09-03 Eric De Giuli

We present a constructive approach to bounded $\ell_2$-gain adaptive control with noisy measurements for linear time-invariant scalar systems with uncertain parameters belonging to a finite set. The gain bound refers to the closed-loop…

最优化与控制 · 数学 2022-02-18 Olle Kjellqvist , Anders Rantzer

In this note, we observe the behavior of gradient flow and discrete and noisy gradient descent in some simple settings. It is commonly noted that addition of noise to gradient descent can affect the trajectory of gradient descent. Here, we…

最优化与控制 · 数学 2019-04-19 Y. Cooper

An efficient policy search algorithm should estimate the local gradient of the objective function, with respect to the policy parameters, from as few trials as possible. Whereas most policy search methods estimate this gradient by observing…

人工智能 · 计算机科学 2012-06-18 Gregory Lawrence , Stuart Russell

The linear quadratic regulator (LQR) problem has reemerged as an important theoretical benchmark for reinforcement learning-based control of complex dynamical systems with continuous state and action spaces. In contrast with nearly all…

机器学习 · 计算机科学 2020-05-04 Benjamin Gravell , Peyman Mohajerin Esfahani , Tyler Summers

While the optimization landscape of policy gradient methods has been recently investigated for partially observed linear systems in terms of both static output feedback and dynamical controllers, they only provide convergence guarantees to…

最优化与控制 · 数学 2023-04-25 Feiran Zhao , Xingyun Fu , Keyou You

In the machine learning literature stochastic gradient descent has recently been widely discussed for its purported implicit regularization properties. Much of the theory, that attempts to clarify the role of noise in stochastic gradient…

机器学习 · 计算机科学 2022-10-21 Alberto Lanconelli , Christopher S. A. Lauria

Can stochastic gradient methods track a moving target? We study the problem of tracking multidimensional time-varying parameters under noisy observations and possible model misspecification. Gradient-based filters update the time-varying…

统计方法学 · 统计学 2026-05-05 Simon Donker van Heel , Rutger-Jan Lange , Bram van Os , Dick van Dijk

Continuous-time Markov decision processes are an important class of models in a wide range of applications, ranging from cyber-physical systems to synthetic biology. A central problem is how to devise a policy to control the system in order…

系统与控制 · 计算机科学 2016-06-01 Ezio Bartocci , Luca Bortolussi , Tomǎš Brázdil , Dimitrios Milios , Guido Sanguinetti

We present a methodology to deploy the stochastic policy gradient method, using actor-critic techniques, when the optimal policy is approximated using a parametric optimization problem, allowing one to enforce safety via hard constraints.…

系统与控制 · 电气工程与系统科学 2024-09-23 Sebastien Gros , Mario Zanon

We consider the problem of adaptive stabilization for discrete-time, multi-dimensional linear systems with bounded control input constraints and unbounded stochastic disturbances, where the parameters of the true system are unknown. To…

系统与控制 · 电气工程与系统科学 2023-04-04 Seth Siriya , Jingge Zhu , Dragan Nešić , Ye Pu

We consider stochastic inviscid dyadic models with energy-preserving noise. It is shown that the models admit weak solutions which are unique in law. Under a certain scaling limit of the noise, the stochastic models converge weakly to a…

概率论 · 数学 2023-05-04 Dejun Luo , Danli Wang

We consider effect of stochastic sources upon self-organization process being initiated with creation of the limit cycle. General expressions obtained are applied to the stochastic Lorenz system to show that departure from equilibrium…

统计力学 · 物理学 2015-05-13 I. A. Shuda , S. S. Borysov , A. I. Olemskoi

Stochastic policies (also known as relaxed controls) are widely used in continuous-time reinforcement learning algorithms. However, executing a stochastic policy and evaluating its performance in a continuous-time environment remain open…

机器学习 · 计算机科学 2025-10-03 Yanwei Jia , Du Ouyang , Yufei Zhang

Robust stability and stochastic stability have separately seen intense study in control theory for many decades. In this work we establish relations between these properties for discrete-time systems and employ them for robust control…

动力系统 · 数学 2020-04-20 Benjamin Gravell , Peyman Mohajerin Esfahani , Tyler Summers

We propose a new analytical method to study stochastic, binary-state models on complex networks. Moving beyond the usual mean-field theories, this alternative approach is based on the introduction of an annealed approximation for…

物理与社会 · 物理学 2016-04-21 Adrián Carro , Raúl Toral , Maxi San Miguel