English
Related papers

Related papers: Permutation Inference under Multi-way Clustering a…

200 papers

We study the clustering problem for mixtures of bounded covariance distributions, under a fine-grained separation assumption. Specifically, given samples from a $k$-component mixture distribution $D = \sum_{i =1}^k w_i P_i$, where each $w_i…

Machine Learning · Computer Science 2023-12-20 Ilias Diakonikolas , Daniel M. Kane , Jasper C. H. Lee , Thanasis Pittas

We introduce a method for validation of results obtained by clustering analysis of data. The method is based on resampling the available data. A figure of merit that measures the stability of clustering solutions against resampling is…

Computational Physics · Physics 2007-05-23 Erel Levine , Eytan Domany

A common problem in machine learning is determining if a variable significantly contributes to a model's prediction performance. This problem is aggravated for datasets, such as gene expression datasets, that suffer the worst case of…

Methodology · Statistics 2023-10-13 Yue Wu , Ted Spaide , Kenji Nakamichi , Russell Van Gelder , Aaron Lee

The topic of this paper is testing exchangeability using e-values in the batch mode, with the Markov model as alternative. The null hypothesis of exchangeability is formalized as a Kolmogorov-type compression model, and the Bayes mixture of…

Methodology · Statistics 2023-05-10 Vladimir Vovk

Transporting findings from a study population to a target population is central to evidence-based decision-making in real-world settings. Most existing methods require individual-level data from both populations to account for covariate…

Methodology · Statistics 2026-03-04 Ying Sheng , Yifei Sun , Chiung-Yu Huang

We consider the problem of inference in shift-share research designs. The choice between existing approaches that allow for unrestricted spatial correlation involves tradeoffs, varying in terms of their validity when there are relatively…

Econometrics · Economics 2022-06-03 Luis Alvarez , Bruno Ferman , Raoni Oliveira

The multimode bunching probability is expected to provide a useful criterion for validating boson sampling experiments. Its applicability, however, is challenged by the existence of anomalous bunching, namely paradoxical situations in which…

Quantum Physics · Physics 2026-01-21 Léo Pioge , Leonardo Novo , Nicolas J. Cerf

We are concerned in clustering continuous data sets subject to non-ignorable missingness. We perform clustering with a specific semi-parametric mixture, under the assumption of conditional independence given the component. The mixture model…

Methodology · Statistics 2021-07-20 Marie Du Roy de Chaumaray , Matthieu Marbac

We propose a test for a change in the mean for a sequence of functional observations that are only partially observed on subsets of the domain, with no information available on the complement. The framework accommodates important scenarios,…

Methodology · Statistics 2025-10-10 Šárka Hudecová , Claudia Kirch

While reliable data-driven decision-making hinges on high-quality labeled data, the acquisition of quality labels often involves laborious human annotations or slow and expensive scientific measurements. Machine learning is becoming an…

Machine Learning · Statistics 2024-03-01 Tijana Zrnic , Emmanuel J. Candès

Many modern estimators require bootstrapping to calculate confidence intervals because either no analytic standard error is available or the distribution of the parameter of interest is non-symmetric. It remains however unclear how to…

Methodology · Statistics 2018-09-13 Michael Schomaker , Christian Heumann

The clustering of bounded data presents unique challenges in statistical analysis due to the constraints imposed on the data values. This paper introduces a novel method for model-based clustering specifically designed for bounded data.…

Methodology · Statistics 2025-05-16 Luca Scrucca

The issue addressed in this paper is that of testing for common breaks across or within equations of a multivariate system. Our framework is very general and allows integrated regressors and trends as well as stationary regressors. The null…

Statistics Theory · Mathematics 2018-01-12 Tatsushi Oka , Pierre Perron

In applications where the study data are collected within cluster units (e.g., patients within transplant centers), it is often of interest to estimate and perform inference on the treatment effects of the cluster units. However, it is…

Applications · Statistics 2023-05-11 Nicholas Hartman , Kevin He

It is conventionally believed that a permutation test should ideally use all permutations. If this is computationally unaffordable, it is believed one should use the largest affordable Monte Carlo sample or (algebraic) subgroup of…

Statistics Theory · Mathematics 2023-11-27 Nick W. Koning

Probabilistic clustering models (or equivalently, mixture models) are basic building blocks in countless statistical models and involve latent random variables over discrete spaces. For these models, posterior inference methods can be…

Machine Learning · Statistics 2020-06-24 Ari Pakman , Yueqi Wang , Catalin Mitelut , JinHyung Lee , Liam Paninski

A clustering outcome for high-dimensional data is typically interpreted via post-processing, involving dimension reduction and subsequent visualization. This destroys the meaning of the data and obfuscates interpretations. We propose…

Machine Learning · Computer Science 2022-09-23 Christian A. Scholbeck , Henri Funk , Giuseppe Casalicchio

Hypothesis tests calibrated by (re)sampling methods (such as permutation, rank and bootstrap tests) are useful tools for statistical analysis, at the computational cost of requiring Monte-Carlo sampling for calibration. It is common and…

Methodology · Statistics 2024-09-30 Ivo V. Stoepker , Rui M. Castro

In scientific studies involving analyses of multivariate data, basic but important questions often arise for the researcher: Is the sample exchangeable, meaning that the joint distribution of the sample is invariant to the ordering of the…

Methodology · Statistics 2023-08-31 Alan J. Aw , Jeffrey P. Spence , Yun S. Song

We propose a novel method for multiple clustering that assumes a co-clustering structure (partitions in both rows and columns of the data matrix) in each view. The new method is applicable to high-dimensional data. It is based on a…