Statistics
Prediction-powered inference (PPI) combines a wall-to-wall prediction map with a small gold-standard sample to give confidence intervals valid whatever the map's quality. Canonical PPI theory starts from i.i.d.\ labelling, whereas spatial…
Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to estimate and rank the causal effects of such interventions. This work focuses on one of…
This paper studies block structured latent variable models, in which observed variables are grouped into distinct blocks based on their relationships with the underlying latent variables. These block structures are prevalent in various…
Observational datasets frequently contain many baseline variables, yet investigators estimating causal effects may not know which variables to include in the adjustment set. Confounding information may also be distributed weakly across many…
Assessing covariate balance across more than two treatment groups has no established omnibus standard: the prevailing practice averages, or takes the maximum of, pairwise standardized mean differences (SMD), while Cohen's f - the classical…
Mortality models that attempt to capture dispersion typically assume a fixed dispersion structure, an assumption that is rarely satisfied in practice and that can lead to miscalibrated uncertainty and poor predictive performance. In this…
Climate change has become a growing concern, particularly in regions experiencing increasingly frequent extreme events. In Bras\'ilia, the capital of Brazil, significant shifts in climate patterns have drawn attention, including episodes of…
This paper proposes a novel extension to Media Mix Modeling (MMM) that introduces a multiplicative interaction between marketing activity and underlying consumer demand. Unlike standard MMM frameworks that assume additive and independent…
Higher-order cumulants capture the non-Gaussian dependence that covariance misses, but they are hard to use in high dimensions. An order-$d$ cumulant tensor has $p^d$ entries, and the plug-in sample cumulant is generally not even…
Metrics like the case-fatality rate and reproduction number are key descriptors of epidemics from the COVID-19 pandemic to the seasonal flu. In retrospect, these quantities enrich our understanding of infectious disease outbreaks; in…
Missing data constitute a pervasive challenge in empirical research. Consequently, there is an ever-growing number of methods designed to address this challenge, with multiple imputation and inverse probability weighting the dominant…
Modern data science increasingly gives rise to hypothesis-testing problems that are not naturally formulated in terms of parameters within prespecified statistical models. One important example is the dynamic evaluation of optimization…
We often observe heterogeneity in longitudinal data, where the mean and variance for certain profiles meaningfully differ from the rest. Some profiles may also exhibit outliers at a limited number of measurements. Using a standard mixed…
Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as…
Uniform stability is a classical tool for controlling the generalization error of a learning algorithm. Bousquet, Klochkov, and Zhivotovskiy (2020) showed that the problem can be reduced to a moment inequality for a sum of weakly…
The Earth Mover's Distance (EMD) is gaining increasing interest among political scientists for assessing similarity in preference distributions. However, there remains a risk of finite-sample upward bias induced by sampling variation in…
Inverse probability weighting (IPW) is widely used to estimate causal effects in observational studies but depends on adequate propensity-score specification. We compare three strategies for addressing treatment assignment heterogeneity:…
Assessing the performance of a basketball team requires the consideration of multiple sources of information. In recent years, the volume and the quality of data generated in sport has increased considerably, particularly in basketball. In…
Introduction: Crash counts on road segments and intersections exhibit differ- ent exposure and connectivity patterns that conventional analyses may obscure. Methodology: A Bayesian negative binomial node edge model was fitted to 8,169 road…
Zero-inflated nonnegative continuous longitudinal data frequently arise in biomedical studies where outcomes consist of a mixture of excess zeros and positive continuous measurements. Two widely used approaches for analyzing such data are…