English
Related papers

Related papers: Optimal Initialization in Depth: Lyapunov Initiali…

200 papers

The Transformer translation model employs residual connection and layer normalization to ease the optimization difficulties caused by its multi-layer encoder/decoder structure. Previous research shows that even with residual connection and…

Computation and Language · Computer Science 2020-05-06 Hongfei Xu , Qiuhui Liu , Josef van Genabith , Deyi Xiong , Jingyi Zhang

Deep learning methods have been widely used in robotic applications, making learning-enabled control design for complex nonlinear systems a promising direction. Although deep reinforcement learning methods have demonstrated impressive…

Systems and Control · Electrical Eng. & Systems 2024-03-19 Zili Wang , Sean B. Andersson , Roberto Tron

We study decentralized optimization over networks where agents cooperatively minimize a smooth (strongly) convex sum of local losses while communicating only with immediate neighbors. Prevailing decentralized methods require either…

Optimization and Control · Mathematics 2026-05-04 Xiaokai Chen , Ilya Kuruzov , Gesualdo Scutari

Activation maximization (AM) strives to generate optimal input stimuli, revealing features that trigger high responses in trained deep neural networks. AM is an important method of explainable AI. We demonstrate that AM fails to produce…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Christoph Linse , Erhardt Barth , Thomas Martinetz

We study the problem of training deep neural networks with Rectified Linear Unit (ReLU) activation function using gradient descent and stochastic gradient descent. In particular, we study the binary classification problem and show that for…

Machine Learning · Computer Science 2018-12-31 Difan Zou , Yuan Cao , Dongruo Zhou , Quanquan Gu

We introduce a variational framework to learn the activation functions of deep neural networks. Our aim is to increase the capacity of the network while controlling an upper-bound of the actual Lipschitz constant of the input-output…

Machine Learning · Computer Science 2023-07-19 Shayan Aziznejad , Harshit Gupta , Joaquim Campos , Michael Unser

The objective of this paper is to enhance the optimization process for neural networks by developing a dynamic learning rate algorithm that effectively integrates exponential decay and advanced anti-overfitting strategies. Our primary…

Machine Learning · Computer Science 2025-08-04 Jatin Chaudhary , Dipak Nidhi , Jukka Heikkonen , Haari Merisaari , Rajiv Kanth

We draw connections between simple neural networks and under-determined linear systems to comprehensively explore several interesting theoretical questions in the study of neural networks. First, we emphatically show that it is unsurprising…

Numerical Analysis · Mathematics 2020-11-02 Austin R. Benson , Anil Damle , Alex Townsend

A delay Lyapunov matrix corresponding to an exponentially stable system of linear time-invariant delay differential equations can be characterized as the solution of a boundary value problem involving a matrix valued delay differential…

Numerical Analysis · Mathematics 2018-08-28 Wim Michiels , Bin Zhou

Overparameterized neural networks often show a benign overfitting property in the sense of achieving excellent generalization behavior despite the number of parameters exceeding the number of training examples. A promising direction to…

Machine Learning · Computer Science 2026-04-23 Yunwen Lei , Yufeng Xie

Nowadays the Lyapunov exponents and Lyapunov dimension have become so widespread and common that they are often used without references to the rigorous definitions or pioneering works. It may lead to a confusion since there are at least two…

Chaotic Dynamics · Physics 2016-03-07 N. V. Kuznetsov , T. A. Alexeeva , G. A. Leonov

We can compare the expressiveness of neural networks that use rectified linear units (ReLUs) by the number of linear regions, which reflect the number of pieces of the piecewise linear functions modeled by such networks. However,…

Machine Learning · Computer Science 2019-12-17 Thiago Serra , Srikumar Ramalingam

This paper introduces the first globally optimal algorithm for the empirical risk minimization problem of two-layer maxout and ReLU networks, i.e., minimizing the number of misclassifications. The algorithm has a worst-case time complexity…

Machine Learning · Computer Science 2026-05-12 Xi He , Yi Miao , Max A. Little

A common method in training neural networks is to initialize all the weights to be independent Gaussian vectors. We observe that by instead initializing the weights into independent pairs, where each pair consists of two identical Gaussian…

Machine Learning · Computer Science 2022-06-28 Alexander Munteanu , Simon Omlor , Zhao Song , David P. Woodruff

Despite the recent success of stochastic gradient descent in deep learning, it is often difficult to train a deep neural network with an inappropriate choice of its initial parameters. Even if training is successful, it has been known that…

Machine Learning · Computer Science 2023-02-10 Cheolhyoung Lee , Kyunghyun Cho

We augment existing generator-side primary frequency control with load-side control that are local, ubiquitous, and continuous. The mechanisms on both the generator and the load sides are decentralized in that their control decisions are…

Systems and Control · Computer Science 2014-03-25 Changhong Zhao , Steven Low

Deep learning has had a far reaching impact in robotics. Specifically, deep reinforcement learning algorithms have been highly effective in synthesizing neural-network controllers for a wide range of tasks. However, despite this empirical…

Robotics · Computer Science 2021-09-30 Hongkai Dai , Benoit Landry , Lujie Yang , Marco Pavone , Russ Tedrake

Neural networks trained via gradient descent with random initialization and without any regularization enjoy good generalization performance in practice despite being highly overparametrized. A promising direction to explain this phenomenon…

Machine Learning · Computer Science 2022-05-17 Hancheng Min , Salma Tarmoun , Rene Vidal , Enrique Mallada

Scaling limits, such as infinite-width limits, serve as promising theoretical tools to study large-scale models. However, it is widely believed that existing infinite-width theory does not faithfully explain the behavior of practical…

Machine Learning · Computer Science 2025-10-28 Moritz Haas , Sebastian Bordt , Ulrike von Luxburg , Leena Chennuru Vankadara

The dependence of the Lyapunov exponent on the closeness parameter, $\epsilon$, in tangent bifurcation systems is investigated. We study and illustrate two averaging procedures for defining Lyapunov exponents in such systems. First, we…

chao-dyn · Physics 2015-06-24 James Hanssen , Walter Wilcox
‹ Prev 1 4 5 6 7 8 10 Next ›