English
Related papers

Related papers: Inference in Regression Discontinuity Designs with…

200 papers

Motivated by modern applications in which one constructs graphical models based on a very large number of features, this paper introduces a new class of cluster-based graphical models, in which variable clustering is applied as an initial…

Machine Learning · Statistics 2020-06-09 Carson Eisenach , Florentina Bunea , Yang Ning , Claudiu Dinicu

Regression discontinuity (RD) is a widely used quasi-experimental design for causal inference. In the standard RD, the assignment to treatment is determined by a continuous pretreatment variable (i.e., running variable) falling above or…

Methodology · Statistics 2020-06-23 Fan Li , Andrea Mercatanti , Taneli Makinen , Andrea Silvestrini

In many applications, data cluster. Failing to take the cluster structure into consideration generally leads to underestimated variances of point estimators and inflated type I errors in hypothesis tests. Many circumstance-dependent…

Methodology · Statistics 2025-07-21 Jiahua Chen , Pengfei Li , Yukun Liu , James V. Zidek

A novel non-parametric estimator of the correlation between grouped measurements of a quantity is proposed in the presence of noise. This work is primarily motivated by functional brain network construction from fMRI data, where brain…

Methodology · Statistics 2023-02-16 Hanâ Lbath , Alexander Petersen , Wendy Meiring , Sophie Achard

This paper develops a general asymptotic theory of local polynomial (LP) regression for spatial data observed at irregularly spaced locations in a sampling region $R_n \subset \mathbb{R}^d$. We adopt a stochastic sampling design that can…

Statistics Theory · Mathematics 2023-12-27 Daisuke Kurisu , Yasumasa Matsuda

It has become standard for empirical studies to conduct inference robust to cluster dependence and heterogeneity. With a small number of clusters, the normal approximation for the $t$-statistics of regression coefficients may be poor. This…

Econometrics · Economics 2026-03-27 Bulat Gafarov , Takuya Ura

Semisupervised methods inevitably invoke some assumption that links the marginal distribution of the features to the regression function of the label. Most commonly, the cluster or manifold assumptions are used which imply that the…

Statistics Theory · Mathematics 2011-12-02 Martin Azizyan , Aarti Singh , Larry Wasserman

We provide finite-sample distribution approximations, that are uniform in the parameter, for inference in linear mixed models. Focus is on variances and covariances of random effects in cases where existing theory fails because their…

Statistics Theory · Mathematics 2025-07-29 Karl Oskar Ekvall , Matteo Bottai

The Latent Block Model (LBM) is a model-based method to cluster simultaneously the $d$ columns and $n$ rows of a data matrix. Parameter estimation in LBM is a difficult and multifaceted problem. Although various estimation strategies have…

Statistics Theory · Mathematics 2020-02-26 Vincent Brault , Christine Keribin , Mahendra Mariadassou

This paper addresses the problem of unsupervised clustering which remains one of the most fundamental challenges in machine learning and artificial intelligence. We propose the clustered generator model for clustering which contains both…

Machine Learning · Statistics 2019-11-20 Dandan Zhu , Tian Han , Linqi Zhou , Xiaokang Yang , Ying Nian Wu

In longitudinal panels and other regression models with unobserved effects, fixed effects estimation is often paired with cluster-robust variance estimation (CRVE) in order to account for heteroskedasticity and un-modeled dependence among…

Methodology · Statistics 2022-11-08 James E. Pustejovsky , Elizabeth Tipton

We consider the problem of detecting whether or not, in a given sensor network, there is a cluster of sensors which exhibit an "unusual behavior." Formally, suppose we are given a set of nodes and attach a random variable to each node. We…

Statistics Theory · Mathematics 2011-03-10 Ery Arias-Castro , Emmanuel J. Candès , Arnaud Durand

Humans are accustomed to environments that contain both regularities and exceptions. For example, at most gas stations, one pays prior to pumping, but the occasional rural station does not accept payment in advance. Likewise, deep neural…

Machine Learning · Computer Science 2021-06-16 Ziheng Jiang , Chiyuan Zhang , Kunal Talwar , Michael C. Mozer

High-dimensional inference methods often rely on coefficient sparsity, an assumption that can be restrictive when signals are dense but individually weak. In such settings, valid inference may still be possible if the covariates exhibit…

Methodology · Statistics 2026-04-14 Wenjun Xiong , Yan Chen , Mingya Long , Qizhai Li

This paper provides a selective review of the statistical network analysis literature focused on clustering and inference problems for stochastic blockmodels and their variants. We survey asymptotic normality results for stochastic…

Statistics Theory · Mathematics 2025-01-24 Joshua Agterberg , Joshua Cape

Regression Discontinuity (RD) designs rely on the continuity of potential outcome means at the cutoff, but this assumption often fails when other treatments or policies are implemented at this cutoff. We characterize the bias in sharp and…

Econometrics · Economics 2025-02-25 Dor Leventer , Daniel Nevo

Predictive mean matching imputation is popular for handling item nonresponse in survey sampling. In this article, we study the asymptotic properties of the predictive mean matching estimator of the population mean. For variance estimation,…

Methodology · Statistics 2018-01-16 Shu Yang , Jae Kwang Kim

We study the econometric properties of so-called donut regression discontinuity (RD) designs, a robustness exercise which involves repeating estimation and inference without the data points in some area around the treatment threshold. This…

Econometrics · Economics 2023-08-29 Cladia Noack , Chistoph Rothe

We re-investigate the asymptotic properties of the traditional OLS (pooled) estimator, $\hat{\beta} _P$, in the context of cluster dependence. The present study considers various scenarios under various restrictions on the cluster sizes and…

Methodology · Statistics 2025-01-31 Subhodeep Dey , Gopal K. Basak , Samarjit Das

The clustering of bounded data presents unique challenges in statistical analysis due to the constraints imposed on the data values. This paper introduces a novel method for model-based clustering specifically designed for bounded data.…

Methodology · Statistics 2025-05-16 Luca Scrucca