English
Related papers

Related papers: Optimal Initialization in Depth: Lyapunov Initiali…

200 papers

"Deep Learning"/"Deep Neural Nets" is a technological marvel that is now increasingly deployed at the cutting-edge of artificial intelligence tasks. This dramatic success of deep learning in the last few years has been hinged on an enormous…

Machine Learning · Computer Science 2021-04-30 Anirbit Mukherjee

We explore the phase diagram of approximation rates for deep neural networks and prove several new theoretical results. In particular, we generalize the existing result on the existence of deep discontinuous phase in ReLU networks to…

Neural and Evolutionary Computing · Computer Science 2021-01-07 Dmitry Yarotsky , Anton Zhevnerchuk

While modern representation learning relies heavily on global error signals, decentralized algorithms driven by local interactions offer a fundamental distributed alternative. However, the macroscopic convergence properties of these…

Machine Learning · Computer Science 2026-04-21 Zilin Li , Weiwei Xu , Xuchun Tong , Xuanbo Lu , Xuanqi Zhao

Autocatalytic chemical networks are dynamical systems whose linearization around zero has a positive Lyapunov exponent; this exponent gives the growth rate of the system in the diluted regime, i.e. for near-zero concentrations. The…

Probability · Mathematics 2025-02-10 Jeremie Unterberger

Understanding and quantifying chaos from data remains challenging. We present a data-driven method for estimating the largest Lyapunov exponent (LLE) from one-dimensional chaotic time series using machine learning. A predictor is trained to…

Chaotic Dynamics · Physics 2025-10-03 A. Velichko , M. Belyaev , P. Boriskov

The rapid increase in the integration of intermittent and stochastic renewable energy resources (RER) introduces challenging issues related to power system stability. Interestingly, identifying grid nodes that can best support stochastic…

Systems and Control · Electrical Eng. & Systems 2025-02-04 Mohamad Kazma , Ahmad F. Taha

Initializing the weights and the biases is a key part of the training process of a neural network. Unlike the subsequent optimization phase, however, the initialization phase has gained only limited attention in the literature. In this…

Machine Learning · Computer Science 2019-09-06 Ingo Steinwart

We study in this paper how to initialize the parameters of multinomial logistic regression (a fully connected layer followed with softmax and cross entropy loss), which is widely used in deep neural network (DNN) models for classification…

Computer Vision and Pattern Recognition · Computer Science 2018-09-18 Bowen Cheng , Rong Xiao , Yandong Guo , Yuxiao Hu , Jianfeng Wang , Lei Zhang

Training neural networks with first order optimisation methods is at the core of the empirical success of deep learning. The scale of initialisation is a crucial factor, as small initialisations are generally associated to a feature…

Machine Learning · Computer Science 2025-09-16 Etienne Boursier , Nicolas Flammarion

Weight initialization remains decisive for neural network optimization, yet existing methods are largely layer-agnostic. We study initialization for deeply-supervised architectures with auxiliary classifiers, where untrained auxiliary heads…

Machine Learning · Computer Science 2026-01-06 Hyunjun Kim

The computation of the entire Lyapunov spectrum for extended dynamical systems is a very time consuming task. If the system is in a chaotic spatio-temporal regime it is possible to approximately reconstruct the Lyapunov spectrum from the…

chao-dyn · Physics 2009-10-31 R. Carretero-González , S. Ørstavik , J. Huke , D. S. Broomhead , J. Stark

This paper presents a small-gain theorem for networks composed of a countably infinite number of finite-dimensional subsystems. Assuming that each subsystem is exponentially input-to-state stable, we show that if the gain operator,…

Optimization and Control · Mathematics 2020-12-02 Christoph Kawan , Andrii Mironchenko , Abdalla Swikir , Navid Noroozi , Majid Zamani

Deep reinforcement learning agents achieve state-of-the-art performance in a wide range of simulated control tasks. However, successful applications to real-world problems remain limited. One reason for this dichotomy is because the learnt…

Machine Learning · Computer Science 2024-11-27 Rory Young , Nicolas Pugeault

Factorized layers--operations parameterized by products of two or more matrices--occur in a variety of deep learning contexts, including compressed model training, certain types of knowledge distillation, and multi-head self-attention…

Machine Learning · Statistics 2022-10-07 Mikhail Khodak , Neil Tenenholtz , Lester Mackey , Nicolò Fusi

Algorithmic stability is a classical framework for analyzing the generalization error of learning algorithms. It predicts that an algorithm has small generalization error if it is insensitive to small perturbations in the training set such…

Machine Learning · Computer Science 2026-02-17 Ouns El Harzli , Yoonsoo Nam , Ilja Kuzborskij , Bernardo Cuenca Grau , Ard A. Louis

Before training a neural net, a classic rule of thumb is to randomly initialize the weights so the variance of activations is preserved across layers. This is traditionally interpreted using the total variance due to randomness in both…

Machine Learning · Computer Science 2019-08-07 Kyle Luther , H. Sebastian Seung

We revisit the initialization of deep residual networks (ResNets) by introducing a novel analytical tool in free probability to the community of deep learning. This tool deals with non-Hermitian random matrices, rather than their…

Machine Learning · Computer Science 2019-02-26 Zenan Ling , Xing He , Robert C. Qiu

We consider dynamical and geometrical aspects of deep learning. For many standard choices of layer maps we display semi-invariant metrics which quantify differences between data or decision functions. This allows us, when considering random…

Machine Learning · Computer Science 2021-04-23 Benny Avelin , Anders Karlsson

Choosing an appropriate learning rate remains a key challenge in scaling depth of modern deep networks. The classical maximal update parameterization ($\mu$P) enforces a fixed per-layer update magnitude, which is well suited to homogeneous…

Machine Learning · Computer Science 2025-12-01 Haosong Zhang , Shenxi Wu , Yichi Zhang , Xi Chen , Wei Lin

We develop a powerful and general method to provide rigorous and accurate upper and lower bounds for Lyapunov exponents of stochastic flows. Our approach is based on computer-assisted tools, the adjoint method and established results on the…

Dynamical Systems · Mathematics 2025-06-02 Maxime Breden , Hugo Chu , Jeroen S. W. Lamb , Martin Rasmussen