English
Related papers

Related papers: Cluster-robust inference with a single treated clu…

200 papers

Many applications from the financial industry successfully leverage clustering algorithms to reveal meaningful patterns among a vast amount of unstructured financial data. However, these algorithms suffer from a lack of interpretability…

Applications · Statistics 2020-07-24 Enguerrand Horel , Kay Giesecke , Victor Storchan , Naren Chittar

An important aspect of multiple hypothesis testing is controlling the significance level, or the level of Type I error. When the test statistics are not independent it can be particularly challenging to deal with this problem, without…

Statistics Theory · Mathematics 2009-03-04 Sandy Clarke , Peter Hall

This paper proposes a selective inference procedure for testing equal predictive ability in panel data settings with unknown heterogeneity. The framework allows predictive performance to vary across unobserved clusters and accounts for the…

Econometrics · Economics 2025-07-29 Oguzhan Akgun , Alain Pirotte , Giovanni Urga , Zhenlin Yang

In cluster-randomized trials, generalized linear mixed models and generalized estimating equations have conventionally been the default analytic methods for estimating the average treatment effect as routine practice. However, recent…

Methodology · Statistics 2025-09-19 Fan Li , Jiaqi Tong , Xi Fang , Chao Cheng , Brennan C. Kahan , Bingkai Wang

A common method to reduce the uncertainty of causal inferences from experiments is to assign treatments in fixed proportions within groups of similar units: blocking. Previous results indicate that one can expect substantial reductions in…

Methodology · Statistics 2015-08-31 Fredrik Sävje

Numerous papers ask how difficult it is to cluster data. We suggest that the more relevant and interesting question is how difficult it is to cluster data sets {\em that can be clustered well}. More generally, despite the ubiquity and the…

Machine Learning · Computer Science 2012-05-23 Amit Daniely , Nati Linial , Michael Saks

Typically clustering algorithms provide clustering solutions with prespecified number of clusters. The lack of a priori knowledge on the true number of underlying clusters in the dataset makes it important to have a metric to compare the…

Machine Learning · Computer Science 2018-11-20 Amber Srivastava , Mayank Baranwal , Srinivasa Salapaka

Many financial and economic variables, including financial returns, exhibit nonlinear dependence, heterogeneity and heavy-tailedness. These properties may make problematic the analysis of (non-)efficiency and volatility clustering in…

Econometrics · Economics 2023-12-01 Rustam Ibragimov , Rasmus Pedersen , Anton Skrobotov

Evaluation of large-scale network systems and applications is usually done in one of three ways: simulations, real deployment on Internet, or on an emulated network testbed such as a cluster. Simulations can study very large systems but…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-12-02 Liang Wang , Jussi Kangasharju

In this paper, we address an issue of finding explainable clusters of class-uniform data in labelled datasets. The issue falls into the domain of interpretable supervised clustering. Unlike traditional clustering, supervised clustering aims…

Machine Learning · Computer Science 2023-07-18 Natallia Kokash , Leonid Makhnist

A data analysis pipeline is a structured sequence of steps that transforms raw data into meaningful insights by integrating multiple analysis algorithms. In many practical applications, analytical findings are obtained only after data pass…

Machine Learning · Statistics 2026-05-04 Yugo Miyata , Tomohiro Shiraishi , Shuichi Nishino , Ichiro Takeuchi

The domain of explainable AI is of interest in all Machine Learning fields, and it is all the more important in clustering, an unsupervised task whose result must be validated by a domain expert. We aim at finding a clustering that has high…

Artificial Intelligence · Computer Science 2024-03-28 Mathieu Guilbert , Christel Vrain , Thi-Bich-Hanh Dao

A new interpoint distance-based measure is proposed to identify the optimal number of clusters present in a data set. Designed in nonparametric approach, it is independent of the distribution of given data. Interpoint distances between the…

Machine Learning · Computer Science 2022-10-18 Soumita Modak

Quality assessments of models in unsupervised learning and clustering verification in particular have been a long-standing problem in the machine learning research. The lack of robust and universally applicable cluster validity scores often…

Machine Learning · Statistics 2018-03-30 Luzie Helfmann , Johannes von Lindheim , Mattes Mollenhauer , Ralf Banisch

Cross-validation plays a fundamental role in Machine Learning, enabling robust evaluation of model performance and preventing overestimation on training and validation data. However, one of its drawbacks is the potential to create data…

Machine Learning · Computer Science 2025-08-28 Afonso Martini Spezia , Thomas Fontanari , Mariana Recamonde-Mendoza

Treatment effect estimands based on win statistics, including the win ratio, win odds, and win difference are increasingly popular targets for summarizing endpoints in clinical trials. Such win estimands offer an intuitive approach for…

Methodology · Statistics 2026-02-13 Kenneth M. Lee , Xi Fang , Fan Li , Michael O. Harhay

There are two notoriously hard problems in cluster analysis, estimating the number of clusters, and checking whether the population to be clustered is not actually homogeneous. Given a dataset, a clustering method and a cluster validation…

Methodology · Statistics 2015-02-10 Christian Hennig , Chien-Ju Lin

We study the canonical fair clustering problem where each cluster is constrained to have close to population-level representation of each group. Despite significant attention, the salient issue of having incomplete knowledge about the group…

Machine Learning · Computer Science 2024-11-21 Sharmila Duppala , Juan Luque , John P. Dickerson , Seyed A. Esmaeili

Evaluation of clinical prediction models across multiple clusters, whether centers or datasets, is becoming increasingly common. A comprehensive evaluation includes an assessment of the agreement between the estimated risks and the observed…

Methodology · Statistics 2026-04-22 Lasai Barreñada , Bavo D. C. Campo , Laure Wynants , Ben Van Calster

Let $(X_{n,i})_{1\le i\le n,n\in\mathbb{N}}$ be a triangular array of row-wise stationary $\mathbb{R}^d$-valued random variables. We use a "blocks method" to define clusters of extreme values: the rows of $(X_{n,i})$ are divided into $m_n$…

Statistics Theory · Mathematics 2020-05-19 Holger Drees , Holger Rootzén