English
Related papers

Related papers: Some issues in robust clustering

200 papers

We study the efficient learnability of high-dimensional Gaussian mixtures in the outlier-robust setting, where a small constant fraction of the data is adversarially corrupted. We resolve the polynomial learnability of this problem when the…

Data Structures and Algorithms · Computer Science 2020-05-14 Ilias Diakonikolas , Samuel B. Hopkins , Daniel Kane , Sushrut Karmalkar

Outlier detection is a core task in data mining with a plethora of algorithms that have enjoyed wide scale usage. Existing algorithms are primarily focused on detection, that is the identification of outliers in a given dataset. In this…

Machine Learning · Computer Science 2019-11-11 Yue Wu , Leman Akoglu , Ian Davidson

In this paper we present methods for exemplar based clustering with outlier selection based on the facility location formulation. Given a distance function and the number of outliers to be found, the methods automatically determine the…

Machine Learning · Computer Science 2014-03-07 Lionel Ott , Linsey Pang , Fabio Ramos , David Howe , Sanjay Chawla

Clustering in high-dimensions poses many statistical challenges. While traditional distance-based clustering methods are computationally feasible, they lack probabilistic interpretation and rely on heuristics for estimation of the number of…

Methodology · Statistics 2023-04-04 Abhinav Natarajan , Maria De Iorio , Andreas Heinecke , Emanuel Mayer , Simon Glenn

We consider the problem of lifting a regular cluster structure on a quasi-affine variety to the ambient affine space and a similar problem of defining a regular pullback of a regular cluster structure under a dominant rational map between…

Commutative Algebra · Mathematics 2026-03-16 Misha Gekhtman , Michael Shapiro , Alek Vainshtein

We consider model-based clustering methods for continuous, correlated data that account for external information available in the presence of mixed-type fixed covariates by proposing the MoEClust suite of models. These models allow…

Methodology · Statistics 2021-07-15 Keefe Murphy , Thomas Brendan Murphy

These lectures cover various aspects of the statistical description of cosmological density fields. Observationally, this consists of the point process defined by galaxies, and the challenge is to relate this to the continuous density field…

Astrophysics · Physics 2007-05-23 J. A. Peacock

A theory of clustering of inertial particles advected by a turbulent velocity field caused by an instability of their spatial distribution is suggested. The reason for the clustering instability is a combined effect of the particles inertia…

Chaotic Dynamics · Physics 2007-05-23 Tov Elperin , Nathan Kleeorin , Victor S. L'vov , Igor Rogachevskii , Dmitry Sokoloff

Clustering provides a common means of identifying structure in complex data, and there is renewed interest in clustering as a tool for the analysis of large data sets in many fields. A natural question is how many clusters are appropriate…

Data Analysis, Statistics and Probability · Physics 2007-05-23 Susanne Still , William Bialek

We present the clustering of galaxy clusters as a useful addition to the common set of cosmological observables. The clustering of clusters probes the large-scale structure of the Universe, extending galaxy clustering analysis to the…

Cosmology and Nongalactic Astrophysics · Physics 2014-02-03 Annalisa Mana , Tommaso Giannantonio , Jochen Weller , Ben Hoyle , Gert Huetsi , Barbara Sartoris

Clustering is a fundamental tool in unsupervised learning, used to group objects by distinguishing between similar and dissimilar features of a given data set. One of the most common clustering algorithms is k-means. Unfortunately, when…

Machine Learning · Statistics 2021-08-17 Olga Dorabiala , J. Nathan Kutz , Aleksandr Aravkin

The aim of the paper is to show that the presence of one possible type of outliers is not connected to that of heavy tails of the distribution. In contrary, typical situation for outliers appearance is the case of compact supported…

Statistics Theory · Mathematics 2018-07-25 Lev B. Klebanov , Irina Volchenkova

We study the problem of outlier robust high-dimensional mean estimation under a finite covariance assumption, and more broadly under finite low-degree moment assumptions. We consider a standard stability condition from the recent robust…

Statistics Theory · Mathematics 2021-03-17 Ilias Diakonikolas , Daniel M. Kane , Ankit Pensia

The increasing needs of clustering massive datasets and the high cost of running clustering algorithms poses difficult problems for users. In this context it is important to determine if a data set is clusterable, that is, it may be…

Machine Learning · Computer Science 2020-01-08 Dan Simovici , Kaixun Hua

The problem of finding groups in data (cluster analysis) has been extensively studied by researchers from the fields of Statistics and Computer Science, among others. However, despite its popularity it is widely recognized that the…

Statistics Theory · Mathematics 2013-10-10 José E. Chacón

The best subset selection (or "best subsets") estimator is a classic tool for sparse regression, and developments in mathematical optimization over the past decade have made it more computationally tractable than ever. Notwithstanding its…

Methodology · Statistics 2022-01-11 Ryan Thompson

We improve current instability-based methods for the selection of the number of clusters $k$ in cluster analysis by developing a normalized cluster instability measure that corrects for the distribution of cluster sizes, a previously…

Machine Learning · Statistics 2018-10-16 Jonas M. B. Haslbeck , Dirk U. Wulff

The percolation properties of clustered networks are analyzed in detail. In the case of weak clustering, we present an analytical approach that allows to find the critical threshold and the size of the giant component. Numerical simulations…

Disordered Systems and Neural Networks · Physics 2009-11-11 M. Angeles Serrano , Marian Boguna

Linear regression is ubiquitous in statistical analysis. It is well understood that conflicting sources of information may contaminate the inference when the classical normality of errors is assumed. The contamination caused by the light…

Methodology · Statistics 2019-06-13 Philippe Gagnon , Alain Desgagné , Mylène Bédard

We provide a complete taxonomic characterization of robust hierarchical clustering methods for directed networks following an axiomatic approach. We begin by introducing three practical properties associated with the notion of robustness in…

Machine Learning · Computer Science 2021-08-21 Gunnar Carlsson , Facundo Mémoli , Santiago Segarra