English
Related papers

Related papers: Compressing Heavy-Tailed Weight Matrices for Non-V…

200 papers

Deep neural networks typically impose significant computational loads and memory consumption. Moreover, the large parameters pose constraints on deploying the model on edge devices such as embedded systems. Tensor decomposition offers a…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Yaping He , Linhao Jiang , Di Wu

Neural networks embedded in safety-sensitive applications such as self-driving cars and wearable health monitors rely on two important techniques: input attribution for hindsight analysis and network compression to reduce its size for…

Machine Learning · Computer Science 2020-10-29 Geondo Park , June Yong Yang , Sung Ju Hwang , Eunho Yang

Finite sample properties of random covariance-type matrices have been the subject of much research. In this paper we focus on the "lower tail" of such a matrix, and prove that it is subgaussian under a simple fourth moment assumption on the…

Probability · Mathematics 2013-12-11 Roberto Imbuzeiro Oliveira

Given an arbitrary continuous probability density function, it is introduced a conjugated probability density, which is defined through the Shannon information associated with its cumulative distribution function. These new densities are…

Statistics Theory · Mathematics 2018-01-26 H. M. de Oliveira , R. J. Cintra

Modern neural networks are highly overparameterized, with capacity to substantially overfit to training data. Nevertheless, these networks often generalize well in practice. It has also been observed that trained networks can often be…

Machine Learning · Statistics 2019-02-26 Wenda Zhou , Victor Veitch , Morgane Austern , Ryan P. Adams , Peter Orbanz

Large-scale deep learning models are well-suited for compression. Across a variety of tasks, methods like pruning, quantization, and knowledge distillation have been used to achieve massive reductions in model parameters with only marginal…

Machine Learning · Computer Science 2026-05-18 Pedram Bakhtiarifard , Tong Chen , Jonathan Wenshøj , Erik B Dam , Raghavendra Selvan

Variational Bayesian Inference is a popular methodology for approximating posterior distributions over Bayesian neural network weights. Recent work developing this class of methods has explored ever richer parameterizations of the…

We study the joint limit distribution of the $k$ largest eigenvalues of a $p\times p$ sample covariance matrix $XX^\T$ based on a large $p\times n$ matrix $X$. The rows of $X$ are given by independent copies of a linear process,…

Probability · Mathematics 2012-10-31 Richard A. Davis , Oliver Pfaffel , Robert Stelzer

Deep neural networks (DNNs) have been successfully applied to many real-world problems, but a complete understanding of their dynamical and computational principles is still lacking. Conventional theoretical frameworks for analysing DNNs…

Machine Learning · Computer Science 2022-03-25 Cheng Kevin Qu , Asem Wardak , Pulin Gong

Both PAC-Bayesian and Sample Compress learning frameworks are instrumental for deriving tight (non-vacuous) generalization bounds for neural networks. We leverage these results in a meta-learning scheme, relying on a hypernetwork that…

Machine Learning · Computer Science 2025-06-06 Benjamin Leblanc , Mathieu Bazinet , Nathaniel D'Amours , Alexandre Drouin , Pascal Germain

In this paper, we investigate and develop a new approach to the numerical analysis and characterization of random fluctuations with heavy-tailed probability distribution function (PDF), such as turbulent heat flow and solar flare…

Statistical Mechanics · Physics 2017-03-22 Mohsen Ghasemi Nezhadhaghighi , Abbas Nakhlband

This paper introduces a new classification scheme - head/tail breaks - in order to find groupings or hierarchy for data with a heavy-tailed distribution. The heavy-tailed distributions are heavily right skewed, with a minority of large…

Data Analysis, Statistics and Probability · Physics 2013-10-22 Bin Jiang

This work studies applications and generalizations of a simple estimation technique that provides exponential concentration under heavy-tailed distributions, assuming only bounded low-order moments. We show that the technique can be used…

Machine Learning · Computer Science 2016-04-19 Daniel Hsu , Sivan Sabato

The presence of non-Gaussian tails is a prevalent characteristic in many financial modeling scenarios, necessitating the use of complex non-Gaussian distributions such as the generalized beta of the second kind (GB2) and the skewed…

Applications · Statistics 2025-12-10 Xing Yan , Yue Zhao , Qi Wu , Wenxuan Ma

We introduce a new class of heavy-tailed distributions for which any weighted average of independent and identically distributed random variables is larger than one such random variable in (usual) stochastic order. We show that many…

Probability · Mathematics 2025-06-18 Yuyu Chen , Seva Shneer

We employ constraints to control the parameter space of deep neural networks throughout training. The use of customized, appropriately designed constraints can reduce the vanishing/exploding gradients problem, improve smoothness of…

Machine Learning · Computer Science 2021-06-22 Benedict Leimkuhler , Tiffany Vlaar , Timothée Pouchon , Amos Storkey

While previous optimization results have suggested that deep neural networks tend to favour low-rank weight matrices, the implications of this inductive bias on generalization bounds remain underexplored. In this paper, we apply Maurer's…

Machine Learning · Computer Science 2024-11-22 Andrea Pinto , Akshay Rangamani , Tomaso Poggio

Recent advances in deep learning have made available large, powerful convolutional neural networks (CNN) with state-of-the-art performance in several real-world applications. Unfortunately, these large-sized models have millions of…

Machine Learning · Computer Science 2020-07-17 Giosuè Cataldo Marinò , Gregorio Ghidoli , Marco Frasca , Dario Malchiodi

Recent theoretical studies have shown that heavy-tails can emerge in stochastic optimization due to `multiplicative noise', even under surprisingly simple settings, such as linear regression with Gaussian data. While these studies have…

Machine Learning · Statistics 2025-05-06 Mert Gurbuzbalaban , Yuanhan Hu , Umut Simsekli , Kun Yuan , Lingjiong Zhu

We provide asymptotic theory for certain functions of the sample autocovariance matrices of a high-dimensional time series with infinite fourth moment. The time series exhibits linear dependence across the coordinates and through time.…

Statistics Theory · Mathematics 2020-01-16 Johannes Heiny , Thomas Mikosch