English
Related papers

Related papers: LDLT L-Lipschitz Network Weight Parameterization I…

200 papers

The LSW theory of Ostwald ripening concerns the time evolution of the size distribution of a dilute system of particles that evolve by diffusional mass transfer with a common mean field. We prove global existence, uniqueness and continuous…

Analysis of PDEs · Mathematics 2007-05-23 Barbara Niethammer , Robert L. Pego

LoRA adapts large language models (LLMs) by restricting updates to low-rank subspaces of pre-trained weights. While this substantially reduces training cost, the effectiveness of adaptation critically depends on which subspace is chosen at…

Machine Learning · Computer Science 2026-05-28 Zhi-Quan Feng , Ying-Jia Lin , Hung-Yu Kao

Training neural network models with discrete (categorical or structured) latent variables can be computationally challenging, due to the need for marginalization over large or combinatorial sets. To circumvent this issue, one typically…

Machine Learning · Computer Science 2020-12-29 Gonçalo M. Correia , Vlad Niculae , Wilker Aziz , André F. T. Martins

We apply new results on free boundary regularity of one-phase almost minimizers in periodic media to obtain a quantitative convergence rate for the shape optimizers of the first Dirichlet eigenvalue in periodic homogenization. We obtain a…

Analysis of PDEs · Mathematics 2022-09-07 William M Feldman

Modeling information spread through a network is one of the key problems of network analysis, with applications in a wide array of areas such as marketing and public health. Most approaches assume that the spread is governed by some…

Social and Information Networks · Computer Science 2025-11-03 Alexander Kagan , Elizaveta Levina , Ji Zhu

We study a family of sparse estimators defined as minimizers of some empirical Lipschitz loss function -- which include the hinge loss, the logistic loss and the quantile regression loss -- with a convex, sparse or group-sparse…

Machine Learning · Statistics 2021-09-23 Antoine Dedieu

We present a novel methodology based on a Taylor expansion of the network output for obtaining analytical expressions for the expected value of the network weights and output under stochastic training. Using these analytical expressions the…

Machine Learning · Statistics 2019-12-19 Anastasia Borovykh

In this article, we study high-dimensional behavior of empirical spectral distributions $\{L_N(t), t\in[0,T]\}$ for a class of $N\times N$ symmetric/Hermitian random matrices, whose entries are generated from the solution of stochastic…

Probability · Mathematics 2020-08-12 Jian Song , Jianfeng Yao , Wangjun Yuan

We analyze the global convergence of gradient descent for deep linear residual networks by proposing a new initialization: zero-asymmetric (ZAS) initialization. It is motivated by avoiding stable manifolds of saddle points. We prove that…

Machine Learning · Computer Science 2019-11-05 Lei Wu , Qingcan Wang , Chao Ma

The logit outputs of a feedforward neural network at initialization are conditionally Gaussian, given a random covariance matrix defined by the penultimate layer. In this work, we study the distribution of this random matrix. Recent work…

Machine Learning · Statistics 2023-06-16 Mufan Bill Li , Mihai Nica , Daniel M. Roy

The minimal norm weight perturbations of DNNs required to achieve a specified change in output are derived and the factors determining its size are discussed. These single-layer exact formulae are contrasted with more generic multi-layer…

Machine Learning · Computer Science 2026-05-19 Bethan Evans , Jared Tanner

We study the implicit bias of Sharpness-Aware Minimization (SAM) when training $L$-layer linear diagonal networks on linearly separable binary classification. For linear models ($L=1$), both $\ell_\infty$- and $\ell_2$-SAM recover the…

Machine Learning · Computer Science 2026-05-19 Chaewon Moon , Dongkuk Si , Chulhee Yun

We establish the fundamental limits in the approximation of Lipschitz functions by deep ReLU neural networks with finite-precision weights. Specifically, three regimes, namely under-, over-, and proper quantization, in terms of minimax…

Machine Learning · Statistics 2024-05-06 Weigutian Ou , Philipp Schenkel , Helmut Bölcskei

It has been noted in existing literature that over-parameterization in ReLU networks generally improves performance. While there could be several factors involved behind this, we prove some desirable theoretical properties at initialization…

Machine Learning · Statistics 2019-10-03 Devansh Arpit , Yoshua Bengio

This paper investigates multilevel initialization strategies for training very deep neural networks with a layer-parallel multigrid solver. The scheme is based on the continuous interpretation of the training problem as a problem of optimal…

Machine Learning · Computer Science 2019-12-20 Eric C. Cyr , Stefanie Günther , Jacob B. Schroder

We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient descent can achieve optimality \emph{without} massive…

Machine Learning · Computer Science 2024-10-11 Semih Cayci , Atilla Eryilmaz

This paper is devoted to the estimation of the Lipschitz constant of general neural network architectures using semidefinite programming. For this purpose, we interpret neural networks as time-varying dynamical systems, where the $k$-th…

Machine Learning · Computer Science 2024-11-26 Patricia Pauli , Dennis Gramlich , Frank Allgöwer

A traditional approach to initialization in deep neural networks (DNNs) is to sample the network weights randomly for preserving the variance of pre-activations. On the other hand, several studies show that during the training process, the…

Machine Learning · Computer Science 2021-02-16 Mert Gurbuzbalaban , Yuanhan Hu

Given a large number of covariates $Z$, we consider the estimation of a high-dimensional parameter $\theta$ in an individualized linear threshold $\theta^T Z$ for a continuous variable $X$, which minimizes the disagreement between…

Statistics Theory · Mathematics 2019-05-28 Huijie Feng , Yang Ning , Jiwei Zhao

Layer-sequential unit-variance (LSUV) initialization - a simple method for weight initialization for deep net learning - is proposed. The method consists of the two steps. First, pre-initialize weights of each convolution or inner-product…

Machine Learning · Computer Science 2016-02-22 Dmytro Mishkin , Jiri Matas