English
Related papers

Related papers: Tail-Aware Information-Theoretic Generalization fo…

200 papers

Given an arbitrary continuous probability density function, it is introduced a conjugated probability density, which is defined through the Shannon information associated with its cumulative distribution function. These new densities are…

Statistics Theory · Mathematics 2018-01-26 H. M. de Oliveira , R. J. Cintra

Large language models (LLMs) are trained on web-scale corpora that exhibit steep power-law distributions, in which the distribution of knowledge is highly long-tailed, with most appearing infrequently. While scaling has improved…

Computation and Language · Computer Science 2026-02-19 Sanket Badhe , Deep Shah , Nehal Kathrotia

An information-theoretic upper bound on the generalization error of supervised learning algorithms is derived. The bound is constructed in terms of the mutual information between each individual training sample and the output of the…

Machine Learning · Computer Science 2020-08-06 Yuheng Bu , Shaofeng Zou , Venugopal V. Veeravalli

In this work, we present a variety of novel information-theoretic generalization bounds for learning algorithms, from the supersample setting of Steinke & Zakynthinou (2020)-the setting of the "conditional mutual information" framework. Our…

Machine Learning · Statistics 2023-06-16 Ziqiao Wang , Yongyi Mao

Long-tailed data is a special type of multi-class imbalanced data with a very large amount of minority/tail classes that have a very significant combined influence. Long-tailed learning aims to build high-performance models on datasets with…

Machine Learning · Computer Science 2024-08-02 Chongsheng Zhang , George Almpanidis , Gaojuan Fan , Binquan Deng , Yanbo Zhang , Ji Liu , Aouaidjia Kamel , Paolo Soda , João Gama

A recent line of empirical studies has demonstrated that SGD might exhibit a heavy-tailed behavior in practical settings, and the heaviness of the tails might correlate with the overall performance. In this paper, we investigate the…

Machine Learning · Computer Science 2023-10-31 Krunoslav Lehman Pavasovic , Alain Durmus , Umut Simsekli

The study of tail behaviour of SGD-induced processes has been attracting a lot of interest, due to offering strong guarantees with respect to individual runs of an algorithm. While many works provide high-probability guarantees, quantifying…

Machine Learning · Computer Science 2026-02-06 Aleksandar Armacki , Dragana Bajović , Dušan Jakovetić , Soummya Kar , Ali H. Sayed

We propose a novel probabilistic model to facilitate the learning of multivariate tail dependence of multiple financial assets. Our method allows one to construct from known random vectors, e.g., standard normal, sophisticated joint…

Risk Management · Quantitative Finance 2020-01-14 Xing Yan , Qi Wu , Wen Zhang

Stochastic volatility processes with heavy-tailed innovations are a well-known model for financial time series. In these models, the extremes of the log returns are mainly driven by the extremes of the i.i.d. innovation sequence which leads…

Probability · Mathematics 2016-03-25 Anja Janssen , Holger Drees

Real-world data are long-tailed, the lack of tail samples leads to a significant limitation in the generalization ability of the model. Although numerous approaches of class re-balancing perform well for moderate class imbalance problems,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Yanbiao Ma , Licheng Jiao , Fang Liu , Shuyuan Yang , Xu Liu , Puhua Chen

It has repeatedly been observed that loss minimization by stochastic gradient descent (SGD) leads to heavy-tailed distributions of neural network parameters. Here, we analyze a continuous diffusion approximation of SGD, called homogenized…

Machine Learning · Statistics 2024-02-05 Zhe Jiao , Martin Keller-Ressel

An influential line of recent work has focused on the generalization properties of unregularized gradient-based learning procedures applied to separable linear classification with exponentially-tailed loss functions. The ability of such…

Machine Learning · Computer Science 2022-06-24 Matan Schliserman , Tomer Koren

Existing alignment methods share a common topology of information flow, where reward information is collected from humans, modeled with preference learning, and used to tune language models. However, this shared topology has not been…

Machine Learning · Computer Science 2025-05-29 Tianyi Qiu , Fanzhi Zeng , Jiaming Ji , Dong Yan , Kaile Wang , Jiayi Zhou , Yang Han , Josef Dai , Xuehai Pan , Yaodong Yang

This paper demonstrates the robustness of Lipschitz-regularized $\alpha$-divergences as objective functionals in generative modeling, showing they enable stable learning across a wide range of target distributions with minimal assumptions.…

Machine Learning · Statistics 2025-09-09 Ziyu Chen , Hyemin Gu , Markos A. Katsoulakis , Luc Rey-Bellet , Wei Zhu

The Weibull tail-coefficient (WTC) plays a crucial role in extreme value statistics when dealing with Weibull-type tails. Several distributions, such as normal, Gamma, Weibull, and Logistic distributions, exhibit this type of tail…

Statistics Theory · Mathematics 2024-02-08 Lígia Henriques-Rodrigues , Frederico Caeiro , M. Ivette Gomes

The imbalance (or long-tail) is the nature of many real-world data distributions, which often induces the undesirable bias of deep classification models toward frequent classes, resulting in poor performance for tail classes. In this paper,…

Machine Learning · Computer Science 2025-10-13 Fudong Lin , Xu Yuan

We suggest a simple Gaussian mixture model for data generation that complies with Feldman's long tail theory (2020). We demonstrate that a linear classifier cannot decrease the generalization error below a certain level in the proposed…

Machine Learning · Computer Science 2023-07-26 Arman Bolatov , Maxat Tezekbayev , Igor Melnykov , Artur Pak , Vassilina Nikoulina , Zhenisbek Assylbekov

Learning the tail behavior of a distribution is a notoriously difficult problem. By definition, the number of samples from the tail is small, and deep generative models, such as normalizing flows, tend to concentrate on learning the body of…

Machine Learning · Computer Science 2022-06-28 Mike Laszkiewicz , Johannes Lederer , Asja Fischer

Data privacy and long-tailed distribution are the norms rather than the exception in many real-world tasks. This paper investigates a federated long-tailed learning (Fed-LT) task in which each client holds a locally heterogeneous dataset;…

Machine Learning · Computer Science 2023-11-28 Zikai Xiao , Zihan Chen , Songshang Liu , Hualiang Wang , Yang Feng , Jin Hao , Joey Tianyi Zhou , Jian Wu , Howard Hao Yang , Zuozhu Liu

This paper introduces a loss-based generalized Bayesian methodology for high-dimensional robust regression with serially correlated errors and predictors. The proposed framework employs a novel scaled pseudo-Huber (SPH) loss function, which…

Methodology · Statistics 2025-03-13 Saptarshi Chakraborty , Kshitij Khare , George Michailidis