English
Related papers

Related papers: Compressing Heavy-Tailed Weight Matrices for Non-V…

200 papers

Statistical learning in high-dimensional spaces is challenging without a strong underlying data structure. Recent advances with foundational models suggest that text and image data contain such hidden structures, which help mitigate the…

Machine Learning · Statistics 2025-02-04 Charles Arnal , Clement Berenfeld , Simon Rosenberg , Vivien Cabannes

The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that with standard bounded gradient variance noise. Most…

Machine Learning · Computer Science 2026-01-28 Hongxu Chen , Ke Wei , Xiaoming Yuan , Luo Luo

This paper investigates deep neural network (DNN) compression from the perspective of compactly representing and storing trained parameters. We explore the previously overlooked opportunity of cross-layer architecture-agnostic…

Computer Vision and Pattern Recognition · Computer Science 2021-11-22 Yuezhou Sun , Wenlong Zhao , Lijun Zhang , Xiao Liu , Hui Guan , Matei Zaharia

Characteristic functions of weighted sums of independent random variables exhibit low-rank structure in the quantized tensor train (QTT) representation, also known as matrix product states (MPS), enabling up to exponential compression of…

Machine Learning · Statistics 2026-03-25 Juan José Rodríguez-Aldavero , Juan José García-Ripoll

Sparse linear regression with ill-conditioned Gaussian random designs is widely believed to exhibit a statistical/computational gap, but there is surprisingly little formal evidence for this belief, even in the form of examples that are…

Data Structures and Algorithms · Computer Science 2022-03-08 Jonathan A. Kelner , Frederic Koehler , Raghu Meka , Dhruv Rohatgi

We investigate a problem estimating coefficients of linear regression under sparsity assumption when covariates and noises are sampled from heavy tailed distributions. Additionally, we consider the situation where not only covariates and…

Machine Learning · Statistics 2024-08-05 Takeyuki Sasai , Hironori Fujisawa

The sample compression theory provides generalization guarantees for predictors that can be fully defined using a subset of the training dataset and a (short) message string, generally defined as a binary sequence. Previous works provided…

Machine Learning · Computer Science 2025-03-12 Mathieu Bazinet , Valentina Zantedeschi , Pascal Germain

We consider a finite collection of independent Hermitian heavy-tailed random matrices of growing dimension. Our model includes the L\'evy matrices proposed by Bouchaud and Cizeau, as well as sparse random matrices with O(1) non-zero entries…

Probability · Mathematics 2024-09-24 Charles Bordenave , Alice Guionnet , Camille Male

Models based on multivariate t distributions are widely applied to analyze data with heavy tails. However, all the marginal distributions of the multivariate t distributions are restricted to have the same degrees of freedom, making these…

Methodology · Statistics 2016-04-08 Zhichao Jiang , Peng Ding

This paper offers a precise analytical characterization of the distribution of returns for a portfolio constituted of assets whose returns are described by an arbitrary joint multivariate distribution. In this goal, we introduce a…

Statistical Mechanics · Physics 2009-10-31 D. Sornette , P. Simonetti , J. V. Andersen

We study the problem of estimating the mean of a distribution in high dimensions when either the samples are adversarially corrupted or the distribution is heavy-tailed. Recent developments in robust statistics have established efficient…

Data Structures and Algorithms · Computer Science 2021-01-20 Samuel B. Hopkins , Jerry Li , Fred Zhang

We present a new framework to measure the intrinsic properties of (deep) neural networks. While we focus on convolutional networks, our framework can be extrapolated to any network architecture. In particular, we evaluate two network…

Machine Learning · Computer Science 2022-05-10 Alberto Badias , Ashis Banerjee

This paper considers the problem of robustly estimating a structured covariance matrix with an elliptical underlying distribution with known mean. In applications where the covariance matrix naturally possesses a certain structure, taking…

Applications · Statistics 2016-06-29 Ying Sun , Prabhu Babu , Daniel P. Palomar

Exploiting the great expressive power of Deep Neural Network architectures, relies on the ability to train them. While current theoretical work provides, mostly, results showing the hardness of this task, empirical evidence usually differs…

Machine Learning · Computer Science 2017-06-05 Shai Shalev-Shwartz , Ohad Shamir , Shaked Shammah

Statistical modeling of high dimensional extremes remains challenging and has generally been limited to moderate dimensions. Understanding structural relationships among variables at their extreme levels is crucial both for constructing…

Methodology · Statistics 2026-01-01 Mihyun Kim , Jeongjin Lee

The components underpinning PLMs -- large weight matrices -- were shown to bear considerable redundancy. Matrix factorization, a well-established technique from matrix theory, has been utilized to reduce the number of parameters in PLM.…

Computation and Language · Computer Science 2023-06-27 Siyu Ren , Kenny Q. Zhu

We obtain a number of new general properties, related to the closedness of the class of long-tailed distributions under convolutions, that are of interest themselves and may be applied in many models that deal with "plus" and/or "max"…

Probability · Mathematics 2015-11-24 Hui Xu , Sergey Foss , Yuebao Wang

This paper considers the problem of robustly estimating the parameters of a heavy-tailed multivariate distribution when the covariance matrix is known to have the structure of a low-rank matrix plus a diagonal matrix as considered in factor…

Computation · Statistics 2019-09-30 Rui Zhou , Junyan Liu , Sandeep Kumar , Daniel P. Palomar

Weight-sharing plays a significant role in the success of many deep neural networks, by increasing memory efficiency and incorporating useful inductive priors about the problem into the network. But understanding how weight-sharing can be…

Machine Learning · Computer Science 2023-12-15 Oscar Chang , Hod Lipson

This paper tackles the problem of robust covariance matrix estimation when the data is incomplete. Classical statistical estimation methodologies are usually built upon the Gaussian assumption, whereas existing robust estimation ones assume…