English
Related papers

Related papers: Head/tail Breaks: A New Classification Scheme for …

200 papers

This paper studies the distributed optimization problem under the influence of heavy-tailed gradient noises. Here, a heavy-tailed noise means that the noise does not necessarily satisfy the bounded variance assumption. Instead, it satisfies…

Optimization and Control · Mathematics 2025-05-12 Chao Sun , Huiming Zhang , Bo Chen , Li Yu

Various molecular interaction networks have been claimed to follow power-law decay for their global connectivity distribution. It has been proposed that there may be underlying generative models that explain this heavy-tailed behavior by…

Molecular Networks · Quantitative Biology 2010-04-20 Adrián López García de Lomana , Qasim K. Beg , G. de Fabritiis , Jordi Villà-Freixa

Convolutional neural networks have achieved great improvement on face recognition in recent years because of its extraordinary ability in learning discriminative features of people with different identities. To train such a well-designed…

Computer Vision and Pattern Recognition · Computer Science 2016-11-29 Xiao Zhang , Zhiyuan Fang , Yandong Wen , Zhifeng Li , Yu Qiao

In this paper, we focus on exploiting the group structure for large-dimensional factor models, which captures the homogeneous effects of common factors on individuals within the same group. In view of the fact that datasets in…

Methodology · Statistics 2024-05-14 Yong He , Xiaoyang Ma , Xingheng Wang , Yalin Wang

In binary classification, imbalance refers to situations in which one class is heavily under-represented. This issue is due to either a data collection process or because one class is indeed rare in a population. Imbalanced classification…

Methodology · Statistics 2022-01-07 Arezou Mojiri , Abbas Khalili , Ali Zeinal Hamadani

Handling multiplicity without losing much power has been a persistent challenge in various fields that often face the necessity of managing numerous statistical tests simultaneously. Recently, $p$-value combination methods based on…

Statistics Theory · Mathematics 2024-02-06 Yeonwoo Rho

We study the empirical version of halfspace depths with the objective of establishing a connection between the rates of convergence and the tail behaviour of the corresponding underlying distributions. The intricate interplay between the…

Statistics Theory · Mathematics 2025-06-03 Sibsankar Singha , Marie Kratz , Sreekar Vadlamani

In the real world, the frequency of occurrence of objects is naturally skewed forming long-tail class distributions, which results in poor performance on the statistically rare classes. A promising solution is to mine tail-class examples to…

Computer Vision and Pattern Recognition · Computer Science 2021-12-16 Gursimran Singh , Lingyang Chu , Lanjun Wang , Jian Pei , Qi Tian , Yong Zhang

Object frequency in the real world often follows a power law, leading to a mismatch between datasets with long-tailed class distributions seen by a machine learning model and our expectation of the model to perform well on all classes. We…

Computer Vision and Pattern Recognition · Computer Science 2020-03-25 Muhammad Abdullah Jamal , Matthew Brown , Ming-Hsuan Yang , Liqiang Wang , Boqing Gong

We propose a semiparametric method for fitting the tail of a heavy-tailed population given a relatively small sample from that population and a larger sample from a related background population. We model the tail of the small sample as an…

Methodology · Statistics 2014-10-21 William Fithian , Stefan Wager

Rank-ordering statistics provides a perspective on the rare, largest elements of a population, whereas the statistics of cumulative distributions are dominated by the more numerous small events. The exponent of a power law distribution can…

Condensed Matter · Physics 2015-06-25 Didier Sornette , Leon Knopoff , Yan Kagan , Christian Vanneste

With the rise of the "big data" phenomenon in recent years, data is coming in many different complex forms. One example of this is multi-way data that come in the form of higher-order tensors such as coloured images and movie clips.…

Methodology · Statistics 2021-06-17 Michael P. B. Gallaugher , Peter A. Tait , Paul D. McNicholas

Modern risk modelling approaches deal with vectors of multiple components. The components could be, for example, returns of financial instruments or losses within an insurance portfolio concerning different lines of business. One of the…

Probability · Mathematics 2021-05-12 Miriam Hägele , Jaakko Lehtomaa

We propose a simple data model inspired from natural data such as text or images, and use it to study the importance of learning features in order to achieve good generalization. Our data model follows a long-tailed distribution in the…

Machine Learning · Computer Science 2023-01-02 Thomas Laurent , James H. von Brecht , Xavier Bresson

This article describes mathematical methods for estimating the top-tail of the wealth distribution and therefrom the share of total wealth that the richest $p$ percent hold, which is an intuitive measure of inequality. As the data base for…

Applications · Statistics 2018-07-11 Christoph Dalitz

Classification data sets with skewed class proportions are called imbalanced. Class imbalance is a problem since most machine learning classification algorithms are built with an assumption of equal representation of all classes in the…

Machine Learning · Computer Science 2022-12-22 Azal Ahmad Khan

Heavy-tailed distributions, prevalent in a lot of real-world applications such as finance, telecommunications, queuing theory, and natural language processing, are challenging to model accurately owing to their slow tail decay. Bernstein…

Performance · Computer Science 2025-10-31 Abdelhakim Ziani , András Horváth , Paolo Ballarini

In distributed machine learning, data is dispatched to multiple machines for processing. Motivated by the fact that similar data points often belong to the same or similar classes, and more generally, classification rules of high accuracy…

Machine Learning · Computer Science 2016-12-16 Travis Dick , Mu Li , Venkata Krishna Pillutla , Colin White , Maria Florina Balcan , Alex Smola

Real-world data usually exhibits a long-tailed distribution,with a few frequent labels and a lot of few-shot labels. The study of institution name normalization is a perfect application case showing this phenomenon. There are many…

Computation and Language · Computer Science 2023-02-21 Jiexing Qi , Shuhao Li , Zhixin Guo , Yusheng Huang , Chenghu Zhou , Weinan Zhang , Xinbing Wang , Zhouhan Lin

In risk management, tail risks are of crucial importance. The quality of a tail model, which is determined by data from an unknown distribution, depends critically on the subset of data used to model the tail. Based on a suitably weighted…

Methodology · Statistics 2021-01-19 Ingo Hoffmann , Christoph J. Börner