English
Related papers

Related papers: Some Impossibility Results for Inference With Clus…

200 papers

Unsupervised clustering of feature matrix data is an indispensible technique for exploratory data analysis and quality control of experimental data. However, clusters are difficult to assess for statistical significance in an objective way.…

Statistics Theory · Mathematics 2021-10-01 James Mathews , Cameron Crowe , Rami Vanguri , Margaret Callahan , Travis Hollmann , Saad Nadeem

As predictive algorithms grow in popularity, using the same dataset to both train and test a new model has become routine across research, policy, and industry. Sample-splitting attains valid inference on model properties by using separate…

Econometrics · Economics 2025-11-27 Bruno Fava

Traditional statistical inference in cluster randomized trials typically invokes the asymptotic theory that requires the number of clusters to approach infinity. In this article, we propose an alternative conformal causal inference…

Methodology · Statistics 2024-10-03 Bingkai Wang , Fan Li , Mengxin Yu

Modern science increasingly relies on ever-growing observational datasets and automated inference pipelines, under the implicit belief that accumulating more data makes scientific conclusions more reliable. Here we show that this belief can…

Machine Learning · Computer Science 2026-02-06 Zhipeng Zhang , Kai Li

It has become standard for empirical studies to conduct inference robust to cluster dependence and heterogeneity. With a small number of clusters, the normal approximation for the $t$-statistics of regression coefficients may be poor. This…

Econometrics · Economics 2026-03-27 Bulat Gafarov , Takuya Ura

This paper considers estimation of a univariate density from an individual numerical sequence. It is assumed that (i) the limiting relative frequencies of the numerical sequence are governed by an unknown density, and (ii) there is a known…

Probability · Mathematics 2008-06-19 Andrew B. Nobel , Gusztav Morvai , Sanjeev R. Kulkarni

Several methods have been proposed to estimate the number of clusters in a dataset; the basic ideal behind all of them has been to study an index that measures inter-cluster separation and intra-cluster cohesion over a range of cluster…

Computer Vision and Pattern Recognition · Computer Science 2016-01-12 Bhaskar Mukhoty , Ruchir Gupta , Y. N. Singh

In spite of recent contributions to the literature, informative cluster size settings are not well known and understood. In this paper, we give a formal definition of the problem and describe it from different viewpoints. Data generating…

Statistics Theory · Mathematics 2018-03-06 Jaakko Nevalainen , Somnath Datta , Hannu Oja

The independence clustering problem is considered in the following formulation: given a set $S$ of random variables, it is required to find the finest partitioning $\{U_1,\dots,U_k\}$ of $S$ into clusters such that the clusters…

Machine Learning · Computer Science 2017-03-21 Daniil Ryabko

A key challenge in causal inference from observational studies is the identification and estimation of causal effects in the presence of unmeasured confounding. In this paper, we introduce a novel approach for causal inference that…

Methodology · Statistics 2022-10-17 Ying Zhou , Dingke Tang , Dehan Kong , Linbo Wang

If the state space of a homogeneous continuous-time Markov chain is too large, making inferences - here limited to determining marginal or limit expectations - becomes computationally infeasible. Fortunately, the state space of such a chain…

Probability · Mathematics 2018-06-01 Alexander Erreygers , Jasper De Bock

Motivated by modern applications in which one constructs graphical models based on a very large number of features, this paper introduces a new class of cluster-based graphical models, in which variable clustering is applied as an initial…

Machine Learning · Statistics 2020-06-09 Carson Eisenach , Florentina Bunea , Yang Ning , Claudiu Dinicu

Various methods have recently been proposed to estimate causal effects with confidence intervals that are uniformly valid over a set of data generating processes when high-dimensional nuisance models are estimated by post-model-selection or…

Methodology · Statistics 2025-10-07 Niloofar Moosavi , Tetiana Gorbach , Xavier de Luna

In this paper, we propose a general framework for combining evidence of varying quality to estimate underlying binary latent variables in the presence of restrictions imposed to respect the scientific context. The resulting algorithms…

Methodology · Statistics 2018-08-28 Zhenke Wu , Livia Casciola-Rosen , Antony Rosen , Scott L. Zeger

This paper examines methods of causal inference based on groupwise matching when we observe multiple large groups of individuals over several periods. We formulate causal inference validity through a generalized matching condition,…

Econometrics · Economics 2026-03-24 Ratzanyel Rincón , Kyungchul Song

Measuring conditional dependencies among the variables of a network is of great interest to many disciplines. This paper studies some shortcomings of the existing dependency measures in detecting direct causal influences or their lack of…

Machine Learning · Statistics 2017-06-05 Jalal Etesami , Kun Zhang , Negar Kiyavash

A new interpoint distance-based measure is proposed to identify the optimal number of clusters present in a data set. Designed in nonparametric approach, it is independent of the distribution of given data. Interpoint distances between the…

Machine Learning · Computer Science 2022-10-18 Soumita Modak

We study the uncertainty in galaxy cluster mass estimates derived from X-ray data assuming hydrostatic equilibrium (HE) for the intra cluster gas. Using a Monte-Carlo procedure we generate a general class of mass models allowing very…

Astrophysics · Physics 2007-05-23 C. Balland , A. Blanchard

Visual clustering is a common perceptual task in scatterplots that supports diverse analytics tasks (e.g., cluster identification). However, even with the same scatterplot, the ways of perceiving clusters (i.e., conducting visual…

Human-Computer Interaction · Computer Science 2023-08-14 Hyeon Jeon , Ghulam Jilani Quadri , Hyunwook Lee , Paul Rosen , Danielle Albers Szafir , Jinwook Seo

Clustering is an underspecified task: there are no universal criteria for what makes a good clustering. This is especially true for relational data, where similarity can be based on the features of individuals, the relationships between…

Machine Learning · Statistics 2017-09-29 Sebastijan Dumancic , Hendrik Blockeel
‹ Prev 1 4 5 6 7 8 10 Next ›