English
Related papers

Related papers: Adjusting the adjusted Rand Index -- A multinomial…

200 papers

Cluster-randomized trials (CRTs) involve randomizing entire groups of participants -- called clusters -- to treatment arms but are often comprised of a limited or fixed number of available clusters. While covariate adjustment can account…

Methodology · Statistics 2022-11-29 Angela Y. Zhu , Nandita Mitra , Karla Hemming , Michael O. Harhay , Fan Li

Pedestrian attribute recognition is an important multi-label classification problem. Although the convolutional neural networks are prominent in learning discriminative features from images, the data imbalance in multi-label setting for…

Computer Vision and Pattern Recognition · Computer Science 2020-05-22 Yang Hu , Xiaying Bai , Pan Zhou , Fanhua Shang , Shengmei Shen

Detecting and classifying abnormal system states is critical for condition monitoring, but supervised methods often fall short due to the rarity of anomalies and the lack of labeled data. Therefore, clustering is often used to group similar…

Machine Learning · Computer Science 2025-01-14 Ferdinand Rewicki , Joachim Denzler , Julia Niebling

Empirical likelihood is a popular nonparametric or semi-parametric statistical method with many nice statistical properties. Yet when the sample size is small, or the dimension of the accompanying estimating function is high, the…

Statistics Theory · Mathematics 2010-10-05 Yukun Liu , Jiahua Chen

$k$-means clustering is a well-studied problem due to its wide applicability. Unfortunately, there exist strong theoretical limits on the performance of any algorithm for the $k$-means problem on worst-case inputs. To overcome this barrier,…

Machine Learning · Computer Science 2022-03-22 Jon C. Ergun , Zhili Feng , Sandeep Silwal , David P. Woodruff , Samson Zhou

Propensity score weighting is a tool for causal inference to adjust for measured confounders. Survey data are often collected under complex sampling designs such as multistage cluster sampling, which presents challenges for propensity score…

Methodology · Statistics 2016-07-27 Shu Yang

Reliable causal effect estimation from observational data requires adjustment for confounding and sufficient overlap in covariate distributions between treatment groups. However, in high-dimensional settings, lack of overlap often inflates…

Methodology · Statistics 2025-03-21 Linying Yang , Robin J. Evans

This paper develops a clustering method that takes advantage of the sturdiness of model-based clustering, while attempting to mitigate some of its pitfalls. First, we note that standard model-based clustering likely leads to the same number…

Machine Learning · Statistics 2022-12-09 Miguel de Carvalho , Gabriel Martos Venturini , Andrej Svetlošák

The purpose of this paper is to analyze the degree index and clustering index in random graphs. The degree index in our setup is a certain measure of degree irregularity whose basic properties are well studied in the literature, and the…

Probability · Mathematics 2026-03-11 Ümit Işlak , Barış Yeşiloğlu

Covariate-adaptive randomization (CAR) procedures are frequently used in comparative studies to increase the covariate balance across treatment groups. However, because randomization inevitably uses the covariate information when forming…

Statistics Theory · Mathematics 2022-07-08 Wei Ma , Yichen Qin , Yang Li , Feifang Hu

Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying…

Machine Learning · Statistics 2008-03-26 Benhuai Xie , Wei Pan , Xiaotong Shen

While K-means is known to be a standard clustering algorithm, its performance may be compromised due to the presence of outliers and high-dimensional noisy variables. This paper proposes adaptively robust and sparse K-means clustering…

Computation · Statistics 2024-11-08 Hao Li , Shonosuke Sugasawa , Shota Katayama

Multireference alignment (MRA) is the problem of estimating a signal from many noisy and cyclically shifted copies of itself. In this paper, we consider an extension called heterogeneous MRA, where $K$ signals must be estimated, and each…

Information Theory · Computer Science 2018-02-02 Nicolas Boumal , Tamir Bendory , Roy R. Lederman , Amit Singer

Sample size determination for cluster randomised trials (CRTs) is challenging as it requires robust estimation of the intra-cluster correlation coefficient (ICC). Typically, the sample size is chosen to provide a certain level of power to…

Applications · Statistics 2023-08-23 S. Faye Williamson , Svetlana V. Tishkovskaya , Kevin J. Wilson

In randomized trials, repeated measures of the outcome are routinely collected. The mixed model for repeated measures (MMRM) leverages the information from these repeated outcome measures, and is often used for the primary analysis to…

Methodology · Statistics 2023-07-20 Bingkai Wang , Yu Du

In this work, we investigate the behavior of ridge regression in an overparameterized binary classification task. We assume examples are drawn from (anisotropic) class-conditional cluster distributions with opposing means and we allow for…

Machine Learning · Statistics 2025-03-12 Alexander Tsigler , Luiz F. O. Chamon , Spencer Frei , Peter L. Bartlett

To increase statistical efficiency in a randomized experiment, researchers often use stratification (i.e., blocking) in the design stage. However, conventional practices of stratification fail to exploit valuable information about the…

Methodology · Statistics 2025-10-28 Zikai Li

This paper proposes a novel, nonparametric, interpoint distance-based measure to investigate whether there exist any groups in a set of given data, and if so then, how many groups are prevailing in total. It is a cluster accuracy index…

Methodology · Statistics 2026-05-21 Soumita Modak

Although exchangeable processes from Bayesian nonparametrics have been used as a generating mechanism for random partition models, we deviate from this paradigm to explicitly incorporate clustering information in the formulation of our…

Methodology · Statistics 2024-10-28 David B. Dahl , Richard L. Warr , Thomas P. Jensen

In this paper, a novel unsupervised low-rank representation model, i.e., Auto-weighted Low-Rank Representation (ALRR), is proposed to construct a more favorable similarity graph (SG) for clustering. In particular, ALRR enhances the…

Machine Learning · Computer Science 2021-04-27 Zhiqiang Fu , Yao Zhao , Dongxia Chang , Xingxing Zhang , Yiming Wang
‹ Prev 1 3 4 5 6 7 10 Next ›