English
Related papers

Related papers: LDLT L-Lipschitz Network Weight Parameterization I…

200 papers

We consider a latent space model for dynamic networks, where our objective is to estimate the pairwise inner products plus the intercept of the latent positions. To balance posterior inference and computational scalability, we consider a…

Machine Learning · Statistics 2024-10-16 Peng Zhao , Anirban Bhattacharya , Debdeep Pati , Bani K. Mallick

We study random one-Lipschitz integer functions $f$ on the vertices of a finite connected graph, sampled according to the weight $W(f) = \prod_{\langle v, w \rangle \in E} \mathbf{c}^{ \mathbb{I} \{ f(v) = f(w) \} }$ where $\mathbf{c} \geq…

Probability · Mathematics 2023-09-27 Alex M. Karrila

Ensembles of neural network weight matrices are studied through the training process for the MNIST classification problem, testing the efficacy of matrix models for representing their distributions, under assumptions of Gaussianity and…

Machine Learning · Computer Science 2025-10-08 Edward Hirst , Sanjaye Ramgoolam

Wasserstein distributionally robust optimization (WDRO) strengthens statistical learning under model uncertainty by minimizing the local worst-case risk within a prescribed ambiguity set. Although WDRO has been extensively studied in…

Machine Learning · Statistics 2025-11-12 Changyu Liu , Yuling Jiao , Junhui Wang , Jian Huang

Network pruning is a promising way to generate light but accurate models and enable their deployment on resource-limited edge devices. However, the current state-of-the-art assumes that the effective sub-network and the other superfluous…

Computer Vision and Pattern Recognition · Computer Science 2022-12-08 Yingchun Wang , Song Guo , Jingcai Guo , Weizhan Zhang , Yida Xu , Jie Zhang , Yi Liu

We consider the problem of estimating the underlying edge probabilities of a time-varying network observed at multiple time points. The probability structure is represented by a time-varying graphon that satisfies temporal H\"older…

Methodology · Statistics 2026-05-11 Jeonghwan Lee , Tianxi Li , Adam J. Rothman

We establish a functional large deviation principle for fully connected multi-layer perceptrons with i.i.d. Gaussian weights (LeCun initialization) and general Lipschitz activation functions, including therefore the popular case of ReLU.…

We develop a new real-variable method for weighted $L^p$ estimates. The method is applied to the study of weighted $W^{1, 2}$ estimates in Lipschitz domains for weak solutions of second-order elliptic systems in divergence form with bounded…

Analysis of PDEs · Mathematics 2020-04-08 Zhongwei Shen

In this work, we propose a dissipativity-based method for Lipschitz constant estimation of 1D convolutional neural networks (CNNs). In particular, we analyze the dissipativity properties of convolutional, pooling, and fully connected layers…

Machine Learning · Computer Science 2023-06-21 Patricia Pauli , Dennis Gramlich , Frank Allgöwer

We study in this paper how to initialize the parameters of multinomial logistic regression (a fully connected layer followed with softmax and cross entropy loss), which is widely used in deep neural network (DNN) models for classification…

Computer Vision and Pattern Recognition · Computer Science 2018-09-18 Bowen Cheng , Rong Xiao , Yandong Guo , Yuxiao Hu , Jianfeng Wang , Lei Zhang

Distribution estimation under local differential privacy (LDP) is a fundamental and challenging task. Significant progresses have been made on categorical data. However, due to different evaluation metrics, these methods do not work well…

Machine Learning · Computer Science 2025-09-25 Puning Zhao , Zhikun Zhang , Bo Sun , Li Shen , Liang Zhang , Shaowei Wang , Zhe Liu

We study local linear convergence of gradient descent for finite-width feedforward networks under the squared empirical loss. Prior work shows that GD can remain confined to a Locally Quasi-Convex Region (LQCR) around initialization, but…

Machine Learning · Statistics 2026-05-29 Agnideep Aich , Ashit Baran Aich , Bruce Wade

Wishart random matrices with a sparse or diluted structure are ubiquitous in the processing of large datasets, with applications in physics, biology and economy. In this work we develop a theory for the eigenvalue fluctuations of diluted…

Disordered Systems and Neural Networks · Physics 2018-03-20 Isaac Pérez Castillo , Fernando L. Metz

The weight initialization and the activation function of deep neural networks have a crucial impact on the performance of the training procedure. An inappropriate selection can lead to the loss of information of the input during forward…

Machine Learning · Statistics 2018-10-09 Soufiane Hayou , Arnaud Doucet , Judith Rousseau

An invariant ensemble of $N\times N$ random matrices can be characterised by a joint distribution for eigenvalues $P(\lambda_1,\cdots,\lambda_N)$. The study of the distribution of linear statistics, i.e. of quantities of the form…

Statistical Mechanics · Physics 2017-09-25 Aurélien Grabsch , Christophe Texier

We study differentially private (DP) stochastic optimization (SO) with loss functions whose worst-case Lipschitz parameter over all data may be extremely large or infinite. To date, the vast majority of work on DP SO assumes that the loss…

Machine Learning · Computer Science 2024-10-01 Andrew Lowy , Meisam Razaviyayn

The marginal correlation between two variables is a measure of their linear dependence. The two original variables need not interact directly, because marginal correlation may arise from the mediation of other variables in the system. The…

Methodology · Statistics 2024-12-17 Bautista Arenaza , Sebastián Risau-Gusman , Inés Samengo

Learning conditional densities and identifying factors that influence the entire distribution are vital tasks in data-driven applications. Conventional approaches work mostly with summary statistics, and are hence inadequate for a…

Methodology · Statistics 2022-09-13 Chengliang Tang , Nathan Lenssen , Ying Wei , Tian Zheng

Upcoming LCLS-II/II-HE operation at repetition rates approaching 1MHz demands on-detector data reduction to manage the resulting data volumes. We present a 2D discrete wavelet transform (DWT) pre-processing algorithm that segments…

When optimizing over-parameterized models, such as deep neural networks, a large set of parameters can achieve zero training error. In such cases, the choice of the optimization algorithm and its respective hyper-parameters introduces…

Machine Learning · Computer Science 2019-12-06 Gauthier Gidel , Francis Bach , Simon Lacoste-Julien