English
Related papers

Related papers: The generalized underlap coefficient with an appli…

200 papers

A commonly used characteristic of statistical dependence of adjacency relations in real networks, the clustering coefficient, evaluates chances that two neighbours of a given vertex are adjacent. An extension is obtained by considering…

Applications · Statistics 2013-04-29 Mindaugas Bloznelis , Valentas Kurauskas

Clustering is widely used in unsupervised learning to find homogeneous groups of observations within a dataset. However, clustering mixed-type data remains a challenge, as few existing approaches are suited for this task. This study…

Machine Learning · Statistics 2025-11-26 Badih Ghattas , Alvaro Sanchez San-Benito

In cluster tomography, we propose measuring the number of clusters $N$ intersected by a line segment of length $\ell$ across a finite sample. As expected, the leading order of $N(\ell)$ scales as $a\ell$, where $a$ depends on microscopic…

Disordered Systems and Neural Networks · Physics 2024-02-13 Helen S. Ansell , Samuel J. Frank , István A. Kovács

Mixture distributions are a workhorse model for multimodal data in information theory, signal processing, and machine learning. Yet even when each component density is simple, the differential entropy of the mixture is notoriously hard to…

Information Theory · Computer Science 2026-02-18 Namyoon Lee

Unsupervised machine learning, and in particular data clustering, is a powerful approach for the analysis of datasets and identification of characteristic features occurring throughout a dataset. It is gaining popularity across scientific…

Mesoscale and Nanoscale Physics · Physics 2021-03-23 Maria El Abbassi , Jan Overbeck , Oliver Braun , Michel Calame , Herre S. J. van der Zant , Mickael L. Perrin

We establish a new model-agnostic optimization framework for out-of-distribution generalization via multicalibration, a criterion that ensures a predictor is calibrated across a family of overlapping groups. Multicalibration is shown to be…

Machine Learning · Computer Science 2024-06-04 Jiayun Wu , Jiashuo Liu , Peng Cui , Zhiwei Steven Wu

A coefficient is introduced that quantifies the extent of separation of a random variable $Y$ relative to a number of variables $\mathbf{X} = (X_1, \dots, X_p)$ by skillfully assessing the sensitivity of the relative effects of the…

Methodology · Statistics 2025-03-27 Sebastian Fuchs , Carsten Limbach , Patrick B. Langthaler

Estimating causal effects under exogeneity hinges on two key assumptions: unconfoundedness and overlap. Researchers often argue that unconfoundedness is more plausible when more covariates are included in the analysis. Less discussed is the…

Statistics Theory · Mathematics 2020-01-06 Alexander D'Amour , Peng Ding , Avi Feller , Lihua Lei , Jasjeet Sekhon

Classification and segmentation are crucial in medical image analysis as they enable accurate diagnosis and disease monitoring. However, current methods often prioritize the mutual learning features and shared model parameters, while…

Computer Vision and Pattern Recognition · Computer Science 2023-08-03 Kai Ren , Ke Zou , Xianjie Liu , Yidi Chen , Xuedong Yuan , Xiaojing Shen , Meng Wang , Huazhu Fu

While clustering is ubiquitously used across science and industry, uncertainty in cluster assignments is rarely quantified with rigorous guarantees. We propose a novel conformal inference framework for clustering that returns confidence…

Methodology · Statistics 2026-04-13 YoonHaeng Hur , Anirban Nath , Genevera Allen

The paper presents a new copula based method for measuring dependence between random variables. Our approach extends the Maximum Mean Discrepancy to the copula of the joint distribution. We prove that this approach has several advantageous…

Machine Learning · Computer Science 2019-08-15 Barnabas Poczos , Zoubin Ghahramani , Jeff Schneider

Generalized linear (GL-) statistics are defined as functionals of an U-quantile process and unify different classes of statistics such as U-statistics and L-statistics. We derive a central limit theorem for GL-statistics of strongly mixing…

Statistics Theory · Mathematics 2014-12-02 Svenja Fischer , Roland Fried , Martin Wendler

In the context of cluster analysis and graph partitioning, many external evaluation measures have been proposed in the literature to compare two partitions of the same set. This makes the task of selecting the most appropriate measure for a…

Machine Learning · Computer Science 2021-02-09 Nejat Arinik , Vincent Labatut , Rosa Figueiredo

Cluster analysis is a popular unsupervised learning tool used in many disciplines to identify heterogeneous sub-populations within a sample. However, validating cluster analysis results and determining the number of clusters in a data set…

Machine Learning · Statistics 2024-04-26 Ali Turfah , Xiaoquan Wen

We introduce a novel end-to-end approach for learning to cluster in the absence of labeled examples. Our clustering objective is based on optimizing normalized cuts, a criterion which measures both intra-cluster similarity as well as…

Machine Learning · Computer Science 2019-10-18 Azade Nazi , Will Hang , Anna Goldie , Sujith Ravi , Azalia Mirhoseini

Properties of data are frequently seen to vary depending on the sampled situations, which usually changes along a time evolution or owing to environmental effects. One way to analyze such data is to find invariances, or representative…

Machine Learning · Statistics 2012-09-26 Satoshi Hara , Takashi Washio

A framework for quantifying dependence between random vectors is introduced. With the notion of a collapsing function, random vectors are summarized by single random variables, called collapsed random variables in the framework. Using this…

Methodology · Statistics 2018-01-12 Marius Hofert , Wayne Oldford , Avinash Prasad , Mu Zhu

Controlled experiments are widely used in many applications to investigate the causal relationship between input factors and experimental outcomes. A completely randomized design is usually used to randomly assign treatment levels to…

Methodology · Statistics 2026-05-12 Yiou Li , Lulu Kang , Xiao Huang

In a standard cluster analysis, such as k-means, in addition to clusters locations and distances between them, it's important to know if they are connected or well separated from each other. The main focus of this paper is discovering the…

Machine Learning · Statistics 2017-05-22 Evgeny Bauman , Konstantin Bauman

The recent high level of interest in weighted complex networks gives rise to a need to develop new measures and to generalize existing ones to take the weights of links into account. Here we focus on various generalizations of the…

Statistical Mechanics · Physics 2013-05-29 J. Saramaki , M. Kivela , J. -P. Onnela , K. Kaski , J. Kertesz