中文
相关论文

相关论文: A comparison of Gap statistic definitions with and…

200 篇论文

In cell biology, statistical analysis means testing the hypothesis that there was no effect. This weak form of hypothesis testing neglects effect size, is universally misinterpreted, and is disastrously prone to error when combined with…

其他定量生物学 · 定量生物学 2025-05-13 Josh L. Morgan

Traditional clustering methods typically focus on either cluster-wise global clustering or point-wise local clustering to reveal the intrinsic structures in unlabeled data. Global clustering optimizes an objective function to explore the…

机器学习 · 计算机科学 2025-02-28 Yuxuan Yan , Na Lu , Difei Mei , Ruofan Yan , Youtian Du

Using a measure of clustering derived from the nearest neighbour distribution and the void probability function we are able to distinguish between regular and clustered structures. With an example we show that regularity is a property of a…

天体物理学 · 物理学 2007-05-23 Martin Kerscher

Clustering algorithms have long been the topic of research, representing the more popular side of unsupervised learning. Since clustering analysis is one of the best ways to find some clarity and structure within raw data, this paper…

机器学习 · 计算机科学 2025-11-25 Naitik Gada

The concept of median/consensus has been widely investigated in order to provide a statistical summary of ranking data, i.e. realizations of a random permutation $\Sigma$ of a finite set, $\{1,\; \ldots,\; n\}$ with $n\geq 1$ say. As it…

机器学习 · 计算机科学 2022-01-21 Morgane Goibert , Stéphan Clémençon , Ekhine Irurozki , Pavlo Mozharovskyi

We improve current instability-based methods for the selection of the number of clusters $k$ in cluster analysis by developing a normalized cluster instability measure that corrects for the distribution of cluster sizes, a previously…

机器学习 · 统计学 2018-10-16 Jonas M. B. Haslbeck , Dirk U. Wulff

The Gini index is a number that attempts to measure how equitably a resource is distributed throughout a population, and is commonly used in economics as a measurement of inequality of wealth or income. The Gini index is often defined as…

组合数学 · 数学 2022-02-01 Grant Kopitzke

We define an entropy based on a chosen governing probability distribution. If a certain kind of measurements follow such a distribution it also gives us a suitable scale to study it with. This scale will appear as a link function that is…

数据分析、统计与概率 · 物理学 2007-10-24 Peter Sunehag

The Online Encyclopedia of Integer Sequences (OEIS) is made up of thousands of numerical sequences considered particularly interesting by some mathematicians. The graphic representation of the frequency with which a number n as a function…

概率论 · 数学 2011-06-02 Nicolas Gauvrit , Jean-Paul Delahaye , Hector Zenil

Typically clustering algorithms provide clustering solutions with prespecified number of clusters. The lack of a priori knowledge on the true number of underlying clusters in the dataset makes it important to have a metric to compare the…

机器学习 · 计算机科学 2018-11-20 Amber Srivastava , Mayank Baranwal , Srinivasa Salapaka

This paper demonstrates that basic statistics (mean, variance) of the logarithm of the variate itself can be used in the calculation of differential entropy among random variables known to be multiples and powers of a common underlying…

信息论 · 计算机科学 2009-01-26 Thomas M. Eccardt

Clustering is a data analysis method for extracting knowledge by discovering groups of data called clusters. Among these methods, state-of-the-art density-based clustering methods have proven to be effective for arbitrary-shaped clusters.…

机器学习 · 计算机科学 2023-10-26 Nabil El Malki , Robin Cugny , Olivier Teste , Franck Ravat

We propose a method for reporting how program evaluations reduce gaps between groups, such as the gender or Black-white gap. We first show that the reduction in disparities between groups can be written as the difference in conditional…

计量经济学 · 经济学 2022-01-19 Paul Goldsmith-Pinkham , Karen Jiang , Zirui Song , Jacob Wallace

The main contribution of this paper is a mathematical definition of statistical sparsity, which is expressed as a limiting property of a sequence of probability distributions. The limit is characterized by an exceedance measure~$H$ and a…

统计方法学 · 统计学 2018-05-24 Peter McCullagh , Nicholas Polson

Statistical matching is a technique for integrating two or more data sets when information available for matching records for individual participants across data sets is incomplete. Statistical matching can be viewed as a missing data…

统计方法学 · 统计学 2015-10-14 Jae-kwang Kim , Emily Berg , Taesung Park

This paper presents a clustering approach that allows for rigorous statistical error control similar to a statistical test. We develop estimators for both the unknown number of clusters and the clusters themselves. The estimators depend on…

统计理论 · 数学 2017-07-13 Michael Vogt , Matthias Schmid

Finding the number of meaningful clusters in an unlabeled dataset is important in many applications. Regularized k-means algorithm is a possible approach frequently used to find the correct number of distinct clusters in datasets. The most…

机器学习 · 计算机科学 2025-05-30 Behzad Kamgar-Parsi , Behrooz Kamgar-Parsi

Determining the number of clusters present in a dataset is an important problem in cluster analysis. Conventional clustering techniques generally assume this parameter to be provided up front. %user supplied. %Recently, robustness of any…

机器学习 · 计算机科学 2020-09-01 Jayasree Saha , Jayanta Mukherjee

Clustering is an unsupervised technique of Data Mining. It means grouping similar objects together and separating the dissimilar ones. Each object in the data set is assigned a class label in the clustering process using a distance measure.…

信息检索 · 计算机科学 2011-10-13 Parul Agarwal , M. Afshar Alam , Ranjit Biswas

We use a measure of clustering derived from the nearest neighbour distribution and the void probability function to distinguish between regular and clustered structures. This measure offers a succinct way to incorporate additional…

天体物理学 · 物理学 2007-05-23 Martin Kerscher