English
Related papers

Related papers: Wilcoxon Rank-Based Tests for Clustered Data with …

200 papers

Platform trials evaluate the efficacy of multiple treatments, allowing for late entry of the experimental arms and enabling efficiency gains by sharing controls. The power of individual treatment-control comparisons in such trials can be…

Methodology · Statistics 2024-07-11 Pavla Krotka , Katharina Hees , Peter Jacko , Dominic Magirr , Martin Posch , Marta Bofill Roig

Kernel-based K-means clustering has gained popularity due to its simplicity and the power of its implicit non-linear representation of the data. A dominant concern is the memory requirement since memory scales as the square of the number of…

Machine Learning · Statistics 2016-12-05 Farhad Pourkamali-Anaraki , Stephen Becker

We introduce a software package, denoted as ORKM, that incorporates the Online Regu larized K-Means Clustering (ORKMC) algorithm for processing online multi/single-view data. The function ORKMeans of the ORKMC utilizes a regularization term…

Applications · Statistics 2025-04-23 Miao Yu , Guangbao Guo

Population size estimation is a major challenge in official statistics, social sciences, and natural sciences. The problem can be tackled by applying capture-recapture methods, which vary depending on the number of sources used,…

Applications · Statistics 2025-01-10 Piotr Chlebicki , Maciej Beręsewicz

Data in the form of ranking lists are frequently encountered, and combining ranking results from different sources can potentially generate a better ranking list and help understand behaviors of the rankers. Of interest here are the rank…

Methodology · Statistics 2020-07-20 Xinran Li , Dingdong Yi , Jun S. Liu

We introduce denoiseR, an R package that provides a unified implementation of several state-of-the-art proposals for regularized low rank matrix estimation, along with automatic selection of the regularization parameters. We also extend…

Applications · Statistics 2018-08-09 Julie Josse , Sylvain Sardy , Stefan Wager

Traditional boxplots are widely used for summarizing and visualizing the distribution of numerical data, yet they exhibit significant limitations when applied to skewed or heavy-tailed distributions, often leading to misclassification of…

Methodology · Statistics 2025-11-24 Mustafa Cavus

The robust rank-order test (Fligner and Policello, 1981) was designed as an improvement of the non-parametric Wilcoxon-Mann-Whitney U-test to be more appropriate when the samples being compared have unequal variance. However, it tends to be…

Methodology · Statistics 2020-09-08 Nirvik Sinha

In this paper we introduce a robust to outliers Wilcoxon change-point testing procedure, for distinguishing between short-range dependent time series with a change in mean at unknown time and stationary long-range dependent time series. We…

Methodology · Statistics 2020-07-07 Carina Gerstenberger

Random binnings generated via recursive binary splits are introduced as a way to detect, measure the strength of, and to display the pattern of association between any two variates, whether one or both are continuous or categorical. This…

Methodology · Statistics 2025-04-30 Chris Salahub , Wayne Oldford

The bipartite graph is a ubiquitous data structure that can model the relationship between two entity types: for instance, users and items, queries and webpages. In this paper, we study the problem of ranking vertices of a bipartite graph,…

Information Retrieval · Computer Science 2017-08-16 Xiangnan He , Ming Gao , Min-Yen Kan , Dingxian Wang

Data clustering is a common unsupervised learning method frequently used in exploratory data analysis. However, identifying relevant structures in unlabeled, high-dimensional data is nontrivial, requiring iterative experimentation with…

Human-Computer Interaction · Computer Science 2018-11-29 Marco Cavallo , Çağatay Demiralp

This work presents a guide for the use of some of the functions of the multiColl package in R for the detection of near-multicollinearity. The main contribution, in comparison to other existing packages in R or other econometric software,…

Computation · Statistics 2021-07-08 R. Salmerón , C. B. García , J. García

The following paper presents nonprobsvy -- an R package for inference based on non-probability samples. The package implements various approaches that can be categorized into three groups: prediction-based approach, inverse probability…

Methodology · Statistics 2025-08-21 Łukasz Chrostowski , Piotr Chlebicki , Maciej Beręsewicz

In this paper we propose a nonparametric graphical test based on optimal matching, for assessing the equality of multiple unknown multivariate probability distributions. Our procedure pools the data from the different classes to create a…

Analyzing time-series cross-sectional (also known as longitudinal or panel) data is an important process across a number of fields, including the social sciences, economics, finance, and medicine. PanelMatch is an R package that implements…

Methodology · Statistics 2025-08-19 Adam Rauh , In Song Kim , Kosuke Imai

We introduce novel wild bootstrap procedures for testing superiority in unpaired two-sample survival data. By combining different classical weighted logrank test we obtain tests with broader power behavior. Right censoring within the data…

Statistics Theory · Mathematics 2018-08-20 Marc Ditzhaus , Markus Pauly

We present fastrerandomize, an R package for fast, scalable rerandomization in experimental design. Rerandomization improves precision by discarding treatment assignments that fail a prespecified covariate-balance criterion, but existing…

Computation · Statistics 2026-01-09 Rebecca Goldstein , Connor T. Jerzak , Aniket Kamat , Fucheng Warren Zhu

Estimating the number of clusters (K) is a critical and often difficult task in cluster analysis. Many methods have been proposed to estimate K, including some top performers using resampling approach. When performing cluster analysis in…

Methodology · Statistics 2019-09-05 Yujia Li , Xiangrui Zeng , Chien-Wei Lin , George Tseng

Clustering is widely used for unsupervised structure discovery, yet it offers limited insight into how reliable each individual assignment is. Diagnostics, such as convergence behavior or objective values, may reflect global quality, but…

Machine Learning · Computer Science 2026-05-15 Aggelos Semoglou , John Pavlopoulos