English
Related papers

Related papers: On discrimination between classes of distribution …

200 papers

Whether an extreme observation is an outlier or not, depends strongly on the corresponding tail behaviour of the underlying distribution. We develop an automatic, data-driven method to identify extreme tail behaviour that deviates from the…

Methodology · Statistics 2019-12-06 Shrijita Bhattacharya , Jan Beirlant

Out-of-distribution (OOD) detection is crucial for deploying robust machine learning models. However, when training data follows a long-tailed distribution, the model's ability to accurately detect OOD samples is significantly compromised,…

Machine Learning · Computer Science 2025-09-26 Shuai Feng , Yuxin Ge , Yuntao Du , Mingcai Chen , Chongjun Wang , Lei Feng

This work is concerned with the limiting spectral distribution of rank-based dependency measures in high dimensions. We provide distribution-free results for multivariate empirical versions of Kendall's $\tau$ and Spearman's $\rho$ in a…

Statistics Theory · Mathematics 2025-08-22 Nina Dörnemann , Michael Fleermann , Johannes Heiny

We consider phase-type scale mixture distributions which correspond to distributions of a product of two independent random variables: a phase-type random variable $Y$ and a nonnegative but otherwise arbitrary random variable $S$ called the…

Probability · Mathematics 2017-05-16 Leonardo Rojas-Nandayapa , Wangyue Xie

We benchmark the robustness of maximum likelihood based uncertainty estimation methods to outliers in training data for regression tasks. Outliers or noisy labels in training data results in degraded performances as well as incorrect…

Machine Learning · Computer Science 2022-02-09 Deebul S. Nair , Nico Hochgeschwender , Miguel A. Olivares-Mendez

The majority of traditional classification ru les minimizing the expected probability of error (0-1 loss) are inappropriate if the class probability distributions are ill-defined or impossible to estimate. We argue that in such cases class…

Machine Learning · Statistics 2018-08-14 Robert P. W. Duin , Elzbieta Pekalska

We propose a method to identify and characterize distribution shifts in classification datasets based on optimal transport. It allows the user to identify the extent to which each class is affected by the shift, and retrieves corresponding…

Machine Learning · Computer Science 2022-08-08 Neha Hulkund , Nicolo Fusi , Jennifer Wortman Vaughan , David Alvarez-Melis

This paper deals with testing the equality of $k$ ($k\ge 2$) distribution functions against possible stochastic ordering among them. Two classes of rank tests are proposed for this testing problem. The statistics of the tests under study…

Statistics Theory · Mathematics 2025-06-03 Nikolay I. Nikolov , Eugenia Stoimenova

We propose a two-sample test for high-dimensional means that requires neither distributional nor correlational assumptions, besides some weak conditions on the moments and tail properties of the elements in the random vectors. This…

Methodology · Statistics 2019-04-17 Kaijie Xue , Fang Yao

A recent popular approach to out-of-distribution (OOD) detection is based on a self-supervised learning technique referred to as contrastive learning. There are two main variants of contrastive learning, namely instance and class…

Machine Learning · Computer Science 2022-11-08 Nawid Keshtmand , Raul Santos-Rodriguez , Jonathan Lawry

In samples from a heavy-tailed distribution a second-order approximation is often use to approximate the tail function. Based on the parameters of the approximation, an optimal sample fraction can be estimated which is then used to estimate…

Statistics Theory · Mathematics 2016-12-15 J. Martin van Zyl

In this work, we give a novel general approach for distribution testing. We describe two techniques: our first technique gives sample-optimal testers, while our second technique gives matching sample lower bounds. As a consequence, we…

Data Structures and Algorithms · Computer Science 2016-05-10 Ilias Diakonikolas , Daniel M. Kane

This paper addresses the problem of Generalized Category Discovery (GCD) under a long-tailed distribution, which involves discovering novel categories in an unlabelled dataset using knowledge from a set of labelled categories. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Bingchen Zhao , Kai Han

Understanding the tail behavior of distributions is crucial in statistical theory. For instance, the tail of a distribution plays a ubiquitous role in extreme value statistics, where it is associated with the likelihood of extreme events.…

Statistics Theory · Mathematics 2024-09-11 Rafael Cabral , Maria de Iorio , Andrea Cremaschi

We use a simple method to derive two concentration bounds on the hypergeometric distribution. Comparison with existing results illustrates the advantage of these bounds across different regimes.

Probability · Mathematics 2025-12-18 Vaisakh Mannalath , Víctor Zapatero , Marcos Curty

We consider heavy-tailed distributions and compare the well-known estimators of the tail index, based on extreme value theory with a comparatively recent estimator based on a different idea.

Probability · Mathematics 2016-08-14 Vygantas Paulauskas , Marijus Vaičiulis

A decision must often be made between heavy-tailed and Gaussian errors for a regression or a time series model, and the t-distribution is frequently used when it is assumed that the errors are heavy-tailed distributed. The performance of…

Computation · Statistics 2015-05-11 J. Martin van Zyl

We study the distribution regression problem assuming the distribution of distributions has a doubling measure larger than one. First, we explore the geometry of any distributions that has doubling measure larger than one and build a small…

Machine Learning · Computer Science 2022-03-02 Ilqar Ramazanli

This paper introduces a statistical test inferring whether a variable allows separating two classes by means of a single critical value. Its test statistic is the prediction error of a nonparametric threshold classifier. While this approach…

Methodology · Statistics 2017-07-17 Fabian Schroeder

For large, real-world inductive learning problems, the number of training examples often must be limited due to the costs associated with procuring, preparing, and storing the training examples and/or the computational costs associated with…

Artificial Intelligence · Computer Science 2011-06-24 F. Provost , G. M. Weiss