English
Related papers

Related papers: A comparison of Gap statistic definitions with and…

200 papers

We propose a novel infection spread model based on a random connection graph which represents connections between $n$ individuals. Infection spreads via connections between individuals and this results in a probabilistic cluster formation…

Information Theory · Computer Science 2022-03-30 Batuhan Arasli , Sennur Ulukus

Statistical decision problems lie at the heart of statistical machine learning. The simplest problems are binary and multiclass classification and class probability estimation. Central to their definition is the choice of loss function,…

Machine Learning · Computer Science 2023-08-21 Robert C. Williamson , Zac Cranko

Genetic data are frequently categorical and have complex dependence structures that are not always well understood. For this reason, clustering and classification based on genetic data, while highly relevant, are challenging statistical…

Methodology · Statistics 2016-06-13 Gabriela Bettella Cybis , Marcio Valk , Silvia Regina Costa Lopes

Traditionally, graph quality metrics focus on readability, but recent studies show the need for metrics which are more specific to the discovery of patterns in graphs. Cluster analysis is a popular task within graph analysis, yet there is…

Data Structures and Algorithms · Computer Science 2019-08-22 Amyra Meidiana , Seok-Hee Hong , Peter Eades , Daniel Keim

We propose a bootstrap procedure for data that may exhibit clustering in two or more dimensions. We use insights from the theory of generalized U-statistics to analyze the large-sample properties of statistics that are sample averages from…

Methodology · Statistics 2017-12-06 Konrad Menzel

The use of dual system estimation (DSE) is heavily used in Census Bureau operations. With DSE methods, it is important to implement methods to infer the population size among those with missing data from one or both data sources. The use of…

Computation · Statistics 2026-05-27 Zhiyuan Lu

With the aim to propose a non parametric hypothesis test, this paper carries out a study on the Matching Error (ME), a comparison index of two partitions obtained from the same data set, using for example two clustering methods. This index…

Methodology · Statistics 2019-07-31 Mathias Bourel , Badih Ghattas , Meliza González

The sparsity parameter for clusters of galaxies is obtained in the context of $\Lambda$-gravity. It is shown that, the theoretical estimated values are within the reported error limits of the measured data. Thus, in the future the sparsity…

General Relativity and Quantum Cosmology · Physics 2020-12-01 A. Amekhyan , S. Sargsyan , A. Stepanian

For each partition of a data set into a given number of parts there is a partition such that every part is as much as possible a good model (an "algorithmic sufficient statistic") for the data in that part. Since this can be done for every…

Machine Learning · Computer Science 2022-10-17 Andrew R. Cohen , Paul M. B. Vitányi

This article proposes a novel variance estimator for within-cluster resampling (WCR) and modified within-cluster resampling (MWCR) - two existing methods for analyzing longitudinal data. WCR is a simple but computationally intensive method,…

Methodology · Statistics 2019-12-02 Daniel Xu , Pamela Shaw , Ian Barnett

We study the distribution of the 'gap time', the first time that a large gap appears, in the spatial birth and death point process on $[0,1]$ in which particles are added uniformly in space at rate $\lambda$ and are removed independently at…

Probability · Mathematics 2025-12-05 Eric Foxall , Clément Soubrier

We consider online monitoring of the network event data to detect local changes in a cluster when the affected data stream distribution shifts from one point process to another with different parameters. Specifically, we are interested in…

Methodology · Statistics 2022-12-26 Rui Zhang , Haoyun Wang , Yao Xie

The $k$-means algorithm is often used in clustering applications but its usage requires a complete data matrix. Missing data, however, is common in many applications. Mainstream approaches to clustering missing data reduce the missing data…

Computation · Statistics 2018-06-07 Jocelyn T. Chi , Eric C. Chi , Richard G. Baraniuk

We introduce the concept of pattern graphs--directed acyclic graphs representing how response patterns are associated. A pattern graph represents an identifying restriction that is nonparametrically identified/saturated and is often a…

Methodology · Statistics 2020-12-04 Yen-Chi Chen

The multivariate hypergeometric distribution describes sampling without replacement from a discrete population of elements divided into multiple categories. Addressing a gap in the literature, we tackle the challenge of estimating discrete…

Machine Learning · Computer Science 2024-06-11 Liam Hodgson , Danilo Bzdok

Recently, varextropy has been introduced as a new dispersion index and a measure of information. In this article, we derive the generating function of extropy and present its infinite series representation. Furthermore, we propose new…

Statistics Theory · Mathematics 2025-12-12 Faranak Goodarzi , Somayeh Ghafouri

Causal effects are often characterized with population summaries. These might provide an incomplete picture when there are heterogeneous treatment effects across subgroups. Since the subgroup structure is typically unknown, it is more…

Methodology · Statistics 2026-04-07 Kwangho Kim , Jisu Kim , Edward H. Kennedy

Intrinsic computation refers to how dynamical systems store, structure, and transform historical and spatial information. By graphing a measure of structural complexity against a measure of randomness, complexity-entropy diagrams display…

Chaotic Dynamics · Physics 2009-11-13 David P. Feldman , Carl S. McTague , James P. Crutchfield

Clustering is an unsupervised learning problem that aims to partition unlabelled data points into groups with similar features. Traditional clustering algorithms provide limited insight into the groups they find as their main focus is…

Machine Learning · Computer Science 2022-10-18 Connor Lawless , Oktay Gunluk

Cluster analysis methods seek to partition a data set into homogeneous subgroups. It is useful in a wide variety of applications, including document processing and modern genetics. Conventional clustering methods are unsupervised, meaning…

Methodology · Statistics 2014-07-11 Eric Bair
‹ Prev 1 8 9 10 Next ›