English
Related papers

Related papers: Detection of Sparse Mixtures: Higher Criticism and…

200 papers

Because of its mathematical tractability, the Gaussian mixture model holds a special place in the literature for clustering and classification. For all its benefits, however, the Gaussian mixture model poses problems when the data is skewed…

Applications · Statistics 2020-11-19 Michael P. B. Gallaugher , Paul D. McNicholas , Volodymyr Melnykov , Xuwen Zhu

This paper introduces constrained mixtures for continuous distributions, characterized by a mixture of distributions where each distribution has a shape similar to the base distribution and disjoint domains. This new concept is used to…

Machine Learning · Statistics 2015-03-29 Conrado S. Miranda , Fernando J. Von Zuben

We study the fundamental task of outlier-robust mean estimation for heavy-tailed distributions in the presence of sparsity. Specifically, given a small number of corrupted samples from a high-dimensional heavy-tailed distribution whose mean…

Data Structures and Algorithms · Computer Science 2022-11-30 Ilias Diakonikolas , Daniel M. Kane , Jasper C. H. Lee , Ankit Pensia

Detecting the emergence of an abrupt change-point is a classic problem in statistics and machine learning. Kernel-based nonparametric statistics have been used for this task which enjoy fewer assumptions on the distributions than the…

Machine Learning · Computer Science 2018-11-14 Shuang Li , Yao Xie , Hanjun Dai , Le Song

We study the problem of graph partitioning, or clustering, in sparse networks with prior information about the clusters. Specifically, we assume that for a fraction $\rho$ of the nodes their true cluster assignments are known in advance.…

Physics and Society · Physics 2010-10-06 Armen E. Allahverdyan , Greg Ver Steeg , Aram Galstyan

We investigate the asymptotic behavior of several variants of the scan statistic applied to empirical distributions, which can be applied to detect the presence of an anomalous interval with any length. Of particular interest is Studentized…

Statistics Theory · Mathematics 2020-03-26 Andrew Ying , Wen-Xin Zhou

The scan statistic is widely used in spatial cluster detection applications of inhomogeneous Poisson processes. However, real data may present substantial departure from the underlying Poisson process. One of the possible departures has to…

Methodology · Statistics 2013-11-19 André L. F. Cançado , Cibele Q. da-Silva , Michel F. da Silva

In this work we consider a problem of multi-label classification, where each instance is associated with some binary vector. Our focus is to find a classifier which minimizes false negative discoveries under constraints. Depending on the…

Statistics Theory · Mathematics 2019-03-29 Evgenii Chzhen

Although much of the focus of statistical works on networks has been on static networks, multiple networks are currently becoming more common among network data sets. Usually, a number of network data sets, which share some form of…

Methodology · Statistics 2018-05-29 Sharmodeep Bhattacharyya , Shirshendu Chatterjee

A statistical description of heavy particles suspended in incompressible rough self-similar flows is developed. It is shown that, differently from smooth flows, particles do not form fractal clusters. They rather distribute inhomogeneously…

Chaotic Dynamics · Physics 2007-05-23 J. Bec , M. Cencini , R. Hillerbrand

In this paper, we study the problem of sparse mean estimation under adversarial corruptions, where the goal is to estimate the $k$-sparse mean of a heavy-tailed distribution from samples contaminated by adversarial noise. Existing methods…

Machine Learning · Computer Science 2025-08-26 Jianhao Ma , Rui Ray Chen , Yinghui He , Salar Fattahi , Wei Hu

As we collect additional samples from a data population for which a known density function estimate may have been previously obtained by a black box method, the increased complexity of the data set may result in the true density being…

Machine Learning · Statistics 2022-10-28 Dat Do , Nhat Ho , XuanLong Nguyen

Analysis of three-way data is becoming ever more prevalent in the literature, especially in the area of clustering and classification. Real data, including real three-way data, are often contaminated by potential outlying observations.…

Estimating a high-dimensional sparse covariance matrix from a limited number of samples is a fundamental problem in contemporary data analysis. Most proposals to date, however, are not robust to outliers or heavy tails. Towards bridging…

Statistics Theory · Mathematics 2020-08-04 John Goes , Gilad Lerman , Boaz Nadler

This work considers clustering nodes of a largely incomplete graph. Under the problem setting, only a small amount of queries about the edges can be made, but the entire graph is not observable. This problem finds applications in…

Machine Learning · Computer Science 2021-10-04 Shahana Ibrahim , Xiao Fu

A new variant of the Compressed Sensing problem is investigated when the number of measurements corrupted by errors is upper bounded by some value l but there are no more restrictions on errors. We prove that in this case it is enough to…

Information Theory · Computer Science 2015-09-25 Grigory Kabatiansky , Cedric Tavernier , Serge Vladuts

Score-based divergences have been widely used in machine learning and statistics applications. Despite their empirical success, a blindness problem has been observed when using these for multi-modal distributions. In this work, we discuss…

Machine Learning · Statistics 2025-11-25 Mingtian Zhang , Oscar Key , Peter Hayes , David Barber , Brooks Paige , François-Xavier Briol

Empirical distributions have their in-sample maxima as natural censoring. We look at the "hidden tail", that is, the part of the distribution in excess of the maximum for a sample size of $n$. Using extreme value theory, we examine the…

Statistical Finance · Quantitative Finance 2020-04-14 Nassim Nicholas Taleb

Obtaining rigorous statistical guarantees for generalization under distribution shift remains an open and active research area. We study a setting we call combinatorial distribution shift, where (a) under the test- and…

Machine Learning · Computer Science 2023-08-01 Max Simchowitz , Abhishek Gupta , Kaiqing Zhang

While several papers have investigated computationally and statistically efficient methods for learning Gaussian mixtures, precise minimax bounds for their statistical performance as well as fundamental limits in high-dimensional settings…

Machine Learning · Statistics 2013-06-11 Martin Azizyan , Aarti Singh , Larry Wasserman