English
Related papers

Related papers: Testing Outlier Detection Algorithms for Identifyi…

200 papers

We study the problem of optimal estimation of the density cluster tree under various assumptions on the underlying density. Building up from the seminal work of Chaudhuri et al. [2014], we formulate a new notion of clustering consistency…

Statistics Theory · Mathematics 2019-12-05 Daren Wang , Xinyang Lu , Alessandro Rinaldo

The random cluster model is used to define an upper bound on a distance measure as a function of the number of data points to be classified and the expected value of the number of classes to form in a hybrid K-means and regression…

Machine Learning · Computer Science 2016-02-12 Robert A. Murphy

Particle identification in gaseous detectors traditionally relies on energy loss measurements (dE/dx); however, uncertainties in total energy deposition limit its resolution. The cluster counting technique (dN/dx) offers an alternative…

Traditional anomaly detection techniques onboard satellites are based on reliable, yet limited, thresholding mechanisms which are designed to monitor univariate signals and trigger recovery actions according to specific European Cooperation…

Systems and Control · Electrical Eng. & Systems 2025-07-16 Riccardo Gallon , Fabian Schiemenz , Alessandra Menicucci , Eberhard Gill

Medical data often exhibit characteristics that make cluster analysis particularly challenging, such as missing values, outliers, and cluster features like skewness. Typically, such data would need to be preprocessed -- by cleaning outliers…

Methodology · Statistics 2025-12-16 Jason Pillay , Cristina Tortora , Antonio Punzo , Andriette Bekker

We consider the problem of clustering noisy high-dimensional data points into a union of low-dimensional subspaces and a set of outliers. The number of subspaces, their dimensions, and their orientations are unknown. A probabilistic…

Information Theory · Computer Science 2013-07-19 Reinhard Heckel , Helmut Bölcskei

One of the applications of center-based clustering algorithms such as K-Means is partitioning data points into K clusters. In some examples, the feature space relates to the underlying problem we are trying to solve, and sometimes we can…

Machine Learning · Computer Science 2020-09-23 Ali Hassani , Amir Iranmanesh , Mahdi Eftekhari , Abbas Salemi

Clustering is an unsupervised technique of Data Mining. It means grouping similar objects together and separating the dissimilar ones. Each object in the data set is assigned a class label in the clustering process using a distance measure.…

Information Retrieval · Computer Science 2011-10-13 Parul Agarwal , M. Afshar Alam , Ranjit Biswas

As language models become more general purpose, increased attention needs to be paid to detecting out-of-distribution (OOD) instances, i.e., those not belonging to any of the distributions seen during training. Existing methods for…

Machine Learning · Computer Science 2024-07-19 Aryan Gulati , Xingjian Dong , Carlos Hurtado , Sarath Shekkizhar , Swabha Swayamdipta , Antonio Ortega

Anomaly detection in distributed systems such as High-Performance Computing (HPC) clusters is vital for early fault detection, performance optimisation, security monitoring, reliability in general but also operational insights. Deep Neural…

Machine Learning · Computer Science 2024-05-14 Franz Kevin Stehle , Wainer Vandelli , Giuseppe Avolio , Felix Zahn , Holger Fröning

Clustering algorithms fundamentally group data points by characteristics to identify patterns. Over the past two decades, researchers have extended these methods to analyze trajectories of humans, animals, and vehicles, studying their…

Machine Learning · Computer Science 2025-12-17 Atieh Rahmani , Mansoor Davoodi , Justin M. Calabrese

In a context of a continuous digitalisation of processes, organisations must deal with the challenge of detecting anomalies that can reveal suspicious activities upon an increasing volume of data. To pursue this goal, audit engagements are…

Computational Engineering, Finance, and Science · Computer Science 2024-05-24 A. Herreros-Martínez , R. Magdalena-Benedicto , J. Vila-Francés , A. J. Serrano-López , S. Pérez-Díaz

The Northern Sky Optical Cluster Survey is a project to create an objective catalog of galaxy clusters over the entire high-galactic-latitude Northern sky, with well understood selection criteria. We use the object catalogs generated from…

Astrophysics · Physics 2016-08-30 R. R. Gal , R. R. DeCarvalho , S. C. Odewahn , S. G. Djorgovski , V. E. Margoniner

This paper tries to present a more unified view of clustering, by identifying the relationships between five different clustering algorithms. Some of the results are not new, but they are presented in a cleaner, simpler and more concise…

Machine Learning · Computer Science 2020-06-11 Bernardo A. Gonzalez-Torres

In this paper, we address the problem of unsupervised video anomaly detection (UVAD). The task aims to detect abnormal events in test video using unlabeled videos as training data. The presence of anomalies in the training data poses a…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Jihun Yi , Sungroh Yoon

Large-scale multi-layer networks with large numbers of nodes, edges, and layers arise across various domains, which poses a great computational challenge for the downstream analysis. In this paper, we develop an efficient randomized…

Computation · Statistics 2025-01-10 Wenqing Su , Xiao Guo , Xiangyu Chang , Ying Yang

The problem of clustering noisy and incompletely observed high-dimensional data points into a union of low-dimensional subspaces and a set of outliers is considered. The number of subspaces, their dimensions, and their orientations are…

Machine Learning · Statistics 2015-08-24 Reinhard Heckel , Helmut Bölcskei

We use a deep neural network to detect and place region-of-interest boxes around ultracold atom clouds in absorption and fluorescence images---with the ability to identify and bound multiple clouds within a single image. The neural network…

Quantum Gases · Physics 2021-08-03 Lucas R. Hofer , Milan Krstajić , Péter Juhász , Anna L. Marchant , Robert P. Smith

How can we discover objects we did not know existed within the large datasets that now abound in astronomy? We present an outlier detection algorithm that we developed, based on an unsupervised Random Forest. We test the algorithm on more…

Astrophysics of Galaxies · Physics 2017-01-11 Dalya Baron , Dovi Poznanski

Kernel methods are popular in clustering due to their generality and discriminating power. However, we show that many kernel clustering criteria have density biases theoretically explaining some practically significant artifacts empirically…

Machine Learning · Statistics 2017-12-11 Dmitrii Marin , Meng Tang , Ismail Ben Ayed , Yuri Boykov