中文
相关论文

相关论文: A comparison of Gap statistic definitions with and…

200 篇论文

We present a dataset of word usage graphs (WUGs), where the existing WUGs for multiple languages are enriched with cluster labels functioning as sense definitions. They are generated from scratch by fine-tuned encoder-decoder language…

计算与语言 · 计算机科学 2024-03-28 Mariia Fedorova , Andrey Kutuzov , Nikolay Arefyev , Dominik Schlechtweg

Many recent works on understanding deep learning try to quantify how much individual data instances influence the optimization and generalization of a model. Such attempts reveal characteristics and importance of individual instances, which…

机器学习 · 计算机科学 2023-03-08 Nohyun Ki , Hoyong Choi , Hye Won Chung

Silhouette coefficient is an established internal clustering evaluation measure that produces a score per data point, assessing the quality of its clustering assignment. To assess the quality of the clustering of the whole dataset, the…

机器学习 · 计算机科学 2024-06-25 John Pavlopoulos , Georgios Vardakas , Aristidis Likas

Clustering is the propensity of nodes that share a common neighbour to be connected. It is ubiquitous in many networks but poses many modelling challenges. Clustering typically manifests itself by a higher than expected frequency of…

动力系统 · 数学 2016-01-07 Martin Ritchie , Luc Berthouze , Istvan Z. Kiss

Designing efficient, effective, and consistent metric clustering algorithms is a significant challenge attracting growing attention. Traditional approaches focus on the stability of cluster centers; unfortunately, this neglects the…

The concept of missing at random is central in the literature on statistical analysis with missing data. In general, inference using incomplete data should be based not only on observed data values but should also take account of the…

统计方法学 · 统计学 2013-06-13 Shaun Seaman , John Galati , Dan Jackson , John Carlin

The clustering coefficient quantifies how well connected are the neighbors of a vertex in a graph. In real networks it decreases with the vertex degree, which has been taken as a signature of the network hierarchical structure. Here we show…

统计力学 · 物理学 2007-05-23 Sara Nadiv Soffer , Alexei Vazquez

Assessing equity in treatment of a subpopulation often involves assigning numerical "scores" to all individuals in the full population such that similar individuals get similar scores; matching via propensity scores or appropriate…

统计方法学 · 统计学 2021-10-18 Mark Tygert

The determination of cluster centers generally depends on the scale that we use to analyze the data to be clustered. Inappropriate scale usually leads to unreasonable cluster centers and thus unreasonable results. In this study, we first…

机器学习 · 统计学 2016-10-20 Xiurui Geng , Hairong Tang

A key issue in cluster analysis is the choice of an appropriate clustering method and the determination of the best number of clusters. Different clusterings are optimal on the same data set according to different criteria, and the choice…

统计方法学 · 统计学 2020-06-24 Serhat Emre Akhanli , Christian Hennig

We study statistics of the gaps in Random Average Process (RAP) on a ring with particles hopping symmetrically, except one tracer particle which could be driven. These particles hop either to the left or to the right by a random fraction…

统计力学 · 物理学 2016-05-09 Julien Cividini , Anupam Kundu , Satya N. Majumdar , David Mukamel

We consider a generalized version of the correlation clustering problem, defined as follows. Given a complete graph $G$ whose edges are labeled with $+$ or $-$, we wish to partition the graph into clusters while trying to avoid errors: $+$…

数据结构与算法 · 计算机科学 2016-05-25 Gregory J. Puleo , Olgica Milenkovic

Logistic regression is an important statistical tool for assessing the probability of an outcome based upon some predictive variables. Standard methods can only deal with precisely known data, however many datasets have uncertainties which…

统计方法学 · 统计学 2022-06-09 Nicholas Gray , Scott Ferson

Score matching is a vital tool for learning the distribution of data with applications across many areas including diffusion processes, energy based modelling, and graphical model estimation. Despite all these applications, little work…

机器学习 · 统计学 2025-06-03 Josh Givens , Song Liu , Henry W J Reeve

Distributed data aggregation is an important task, allowing the decentralized determination of meaningful global properties, that can then be used to direct the execution of other applications. The resulting values result from the…

分布式、并行与集群计算 · 计算机科学 2011-10-05 Paulo Jesus , Carlos Baquero , Paulo Sérgio Almeida

Recently, scholars from law and political science have introduced metrics which use only election outcomes (and not district geometry) to assess the presence of partisan gerrymandering. The most high-profile example of such a tool is the…

物理与社会 · 物理学 2018-03-15 Ellen Veomett

Selective clustering annotated using modes of projections (SCAMP) is a new clustering algorithm for data in $\mathbb{R}^p$. SCAMP is motivated from the point of view of non-parametric mixture modeling. Rather than maximizing a…

机器学习 · 统计学 2018-07-30 Evan Greene , Greg Finak , Raphael Gottardo

Persistence is considered in diffusion--limited cluster--cluster aggregation, in one dimension and when the diffusion coefficient of a cluster depends on its size $s$ as $D(s) \sim s^\gamma$. The empty and filled site persistences are…

统计力学 · 物理学 2016-08-16 E. K. O. Hellén , M. J. Alava

Node role explainability in complex networks is very difficult, yet is crucial in different application domains such as social science, neurosciences or computer science. Many efforts have been made on the quantification of hubs revealing…

神经元与认知 · 定量生物学 2023-02-01 Lucrezia Carboni , Michel Dojat , Sophie Achard

The global clustering coefficient serves as a powerful metric for the structural analysis and comparison of complex networks. Random geometric graphs offer a realistic framework for representing the spatial constraints and geometry often…

统计理论 · 数学 2026-02-23 Mingao Yuan , Md. Niamul Islam Sium