中文
相关论文

相关论文: A comparison of Gap statistic definitions with and…

200 篇论文

In many modern applications, there is interest in analyzing enormous data sets that cannot be easily moved across computers or loaded into memory on a single computer. In such settings, it is very common to be interested in clustering.…

统计计算 · 统计学 2020-05-15 Hanyu Song , Yingjian Wang , David B. Dunson

Entropy is a measure of heterogeneity widely used in applied sciences, often when data are collected over space. Recently, a number of approaches has been proposed to include spatial information in entropy. The aim of entropy is to…

统计理论 · 数学 2019-11-12 Linda Altieri , Daniela Cocchi , Giulia Roli

This paper introduces a statistical test inferring whether a variable allows separating two classes by means of a single critical value. Its test statistic is the prediction error of a nonparametric threshold classifier. While this approach…

统计方法学 · 统计学 2017-07-17 Fabian Schroeder

Functional data clustering is to identify heterogeneous morphological patterns in the continuous functions underlying the discrete measurements/observations. Application of functional data clustering has appeared in many publications across…

统计方法学 · 统计学 2022-10-04 Mimi Zhang , Andrew Parnell

Clustering is an unsupervised machine learning methodology where unlabeled elements/objects are grouped together aiming to the construction of well-established clusters that their elements are classified according to their similarity. The…

机器学习 · 统计学 2023-10-20 Dimitrios Saligkaras , Vasileios E. Papageorgiou

This paper clarifies a fundamental difference between causal inference and traditional statistical inference by formalizing a mathematical distinction between their respective parameters. We connect two major approaches to causal inference,…

统计方法学 · 统计学 2025-08-29 Muye Liu , Jun Xie

To assess the presence of gerrymandering, one can consider the shapes of districts or the distribution of votes. The "efficiency gap," which does the latter, plays a central role in a 2016 federal court case on the constitutionality of…

应用统计 · 统计学 2017-05-29 Gregory S. Warrington

The clustering coefficient is a valuable tool for understanding the structure of complex networks. It is widely used to analyze social networks, biological networks, and other complex systems. While there is generally a single common…

物理与社会 · 物理学 2024-01-09 Alexander I Nesterov

Contextuality means non-existence of a joint distribution for random variables recorded under mutually incompatible conditions, subject to certain constraints imposed on how the identity of these variables may change across these…

量子物理 · 物理学 2015-02-06 Ehtibar N. Dzhafarov , Janne V. Kujala

Introduction The tau statistic is a recent second-order correlation function that can assess the magnitude and range of global spatiotemporal clustering from epidemiological data containing geolocations of individual cases and, usually,…

Inference is the process of using facts we know to learn about facts we do not know. A theory of inference gives assumptions necessary to get from the former to the latter, along with a definition for and summary of the resulting…

机器学习 · 统计学 2021-09-27 Beau Coker , Cynthia Rudin , Gary King

We introduce a novel end-to-end approach for learning to cluster in the absence of labeled examples. Our clustering objective is based on optimizing normalized cuts, a criterion which measures both intra-cluster similarity as well as…

机器学习 · 计算机科学 2019-10-18 Azade Nazi , Will Hang , Anna Goldie , Sujith Ravi , Azalia Mirhoseini

Maximum likelihood estimates (MLEs) are asymptotically normally distributed, and this property is used in meta-analyses to test the heterogeneity of estimates, either for a single cluster or for several sub-groups. More recently, MLEs for…

统计理论 · 数学 2022-02-28 Anthony J. Webster

Clustering is one of the fundamental tasks in data analytics and machine learning. In many situations, different clusterings of the same data set become relevant. For example, different algorithms for the same clustering task may return…

最优化与控制 · 数学 2020-04-06 Steffen Borgwardt , Charles Viss

We consider the $1d$ one-component plasma (OCP) in thermal equilibrium, consisting of $N$ equally charged particles on a line, with pairwise Coulomb repulsion and confined by an external harmonic potential. We study two observables: (i) the…

统计力学 · 物理学 2022-06-10 Ana Flack , Satya N. Majumdar , Gregory Schehr

Clustering is a common technique for statistical data analysis, Clustering is the process of grouping the data into classes or clusters so that objects within a cluster have high similarity in comparison to one another, but are very…

机器学习 · 计算机科学 2012-03-12 T Soni Madhulatha

Data clustering is an approach to seek for structure in sets of complex data, i.e., sets of "objects". The main objective is to identify groups of objects which are similar to each other, e.g., for classification. Here, an introduction to…

数据分析、统计与概率 · 物理学 2016-02-17 Alexander K. Hartmann

In many modern statistical problems, the limited available data must be used both to develop the hypotheses to test, and to test these hypotheses-that is, both for exploratory and confirmatory data analysis. Reusing the same dataset for…

统计方法学 · 统计学 2023-07-24 Youngjoo Yun , Rina Foygel Barber

Neutrosophic Statistics means statistical analysis of population or sample that has indeterminate (imprecise, ambiguous, vague, incomplete, unknown) data. For example, the population or sample size might not be exactly determinate because…

人工智能 · 计算机科学 2014-06-10 Florentin Smarandache

In this paper several examples of gaps (lacunes) between dimensions of maximal and submaximal symmetric models are considered, which include investigation of number of independent linear and quadratic integrals of metrics and counting the…

微分几何 · 数学 2012-03-06 Boris Kruglikov