English
Related papers

Related papers: JAWS: Auditing Predictive Uncertainty Under Covari…

200 papers

Quantitative research in the social and behavioral sciences relies heavily on nonlinear posterior functionals such as indirect effects, standardized coefficients, effect sizes, intraclass correlations, and multilevel variance-explained…

Methodology · Statistics 2026-04-07 Nanyu Luo , Feng Ji

We propose the so-called jackknife empirical likelihood approach for the survey data of general unequal probability sampling designs, and analyze parameters defined according to U-statistics. We prove theoretically that jackknife…

Methodology · Statistics 2023-03-28 Mengdong Shang , Xia Chen

This paper studies distribution-free inference in settings where the data set has a hierarchical structure -- for example, groups of observations, or repeated measurements. In such settings, standard notions of exchangeability may not hold.…

Statistics Theory · Mathematics 2025-08-05 Yonghoon Lee , Rina Foygel Barber , Rebecca Willett

The challenge of Out-of-Distribution (OOD) generalization poses a foundational concern for the application of machine learning algorithms to risk-sensitive areas. Inspired by traditional importance weighting and propensity weighting…

Machine Learning · Computer Science 2025-02-12 Han Yu , Yue He , Renzhe Xu , Dongbai Li , Jiayin Zhang , Wenchao Zou , Peng Cui

Covariate shift, a widely used assumption in tackling {\it distributional shift} (when training and test distributions differ), focuses on scenarios where the distribution of the labels conditioned on the feature vector is the same, but the…

Machine Learning · Computer Science 2025-02-24 Deeksha Adil , Jarosław Błasiok

We study the implications of including many covariates in a first-step estimate entering a two-step estimation procedure. We find that a first order bias emerges when the number of \textit{included} covariates is "large" relative to the…

Econometrics · Economics 2018-07-27 Matias D. Cattaneo , Michael Jansson , Xinwei Ma

Missing data are a concern in many real world data sets and imputation methods are often needed to estimate the values of missing data, but data sets with excessive missingness and high dimensionality challenge most approaches to…

Machine Learning · Statistics 2021-04-22 Andrew J. Becker , James P. Bagrow

The Infinitesimal Jackknife is a general method for estimating variances of parametric models, and more recently also for some ensemble methods. In this paper we extend the Infinitesimal Jackknife to estimate the covariance between any two…

Machine Learning · Statistics 2022-09-02 Indrayudh Ghosal , Yunzhe Zhou , Giles Hooker

Bivariate extreme-value distributions have been used in modeling extremes in environmental sciences and risk management. An important issue is estimating the dependence function, such as the Pickands dependence function. Some estimators for…

Statistics Theory · Mathematics 2013-03-21 Liang Peng , Linyi Qian , Jingping Yang

The error or variability of machine learning algorithms is often assessed by repeatedly re-fitting a model with different weighted versions of the observed data. The ubiquitous tools of cross-validation (CV) and the bootstrap are examples…

Methodology · Statistics 2020-02-10 Ryan Giordano , Will Stephenson , Runjing Liu , Michael I. Jordan , Tamara Broderick

We propose a framework, the Neyman Jackknife, for conservative variance estimation in finite-population causal inference under interference. Our approach provides a general, flexible blueprint that enables conservative variance estimation…

Methodology · Statistics 2026-04-28 Bryan Park , Stefan Wager

In the present article, we discuss jackknife empirical likelihood (JEL) and adjusted jackknife empirical likelihood (AJEL) based inference for finding confidence intervals for probability weighted moment (PWM). We obtain the asymptotic…

Methodology · Statistics 2018-07-13 Deepesh Bhati , Sudheesh K Kattumannil , N Sreelakshmi

Conformal prediction methods enjoy strong theoretical and empirical predictive inference performance, provided the data is exchangeable, and predictors are trained in a memoryless fashion. However, these assumptions and constraints are…

Machine Learning · Statistics 2026-05-29 Hanyang Jiang , Rina Foygel Barber , Ashwin Pananjady , Yao Xie

Jackknife instrumental variable estimation (JIVE) is a classic method to leverage many weak instrumental variables (IVs) to estimate linear structural models, overcoming the bias of standard methods like two-stage least squares. In this…

Statistics Theory · Mathematics 2024-10-08 Aurélien Bibaut , Nathan Kallus , Apoorva Lal

For linear regression models with cross-section or panel data, it is natural to assume that the disturbances are clustered in two dimensions. However, the finite-sample properties of two-way cluster-robust tests and confidence intervals are…

Econometrics · Economics 2026-03-13 James G. MacKinnon , Morten Ørregaard Nielsen , Matthew D. Webb

We give an analytical interpretation of how subsample-based internal covariance estimators lead to biased estimates of the covariance, due to underestimating the super-sample covariance (SSC). This includes the jackknife and bootstrap…

Cosmology and Nongalactic Astrophysics · Physics 2018-04-16 Fabien Lacasa , Martin Kunz

Importance weighting (IW) is a golden solver for joint distribution shift, where the joint distributions differ between the training and test data. To solve this problem, IW estimates test-to-training density ratios as importance weights…

Machine Learning · Computer Science 2026-05-26 Tongtong Fang , Nan Lu , Gang Niu , Kenji Fukumizu , Masashi Sugiyama

In a supervised learning problem, given a predicted value that is the output of some trained model, how can we quantify our uncertainty around this prediction? Distribution-free predictive inference aims to construct prediction intervals…

Statistics Theory · Mathematics 2025-07-01 Ruiting Liang , Rina Foygel Barber

Resampling techniques have become increasingly popular for estimation of uncertainty in data collected via surveys. Survey data are also frequently subject to missing data which are often imputed. This note addresses the issue of using…

Methodology · Statistics 2023-11-27 Michael W. Robbins , Lane Burgette , Sebastian Bauhoff

Tabular synthetic data generators are typically trained to match observational distributions, which can yield high conventional utility (e.g., column correlations, predictive accuracy) yet poor preservation of structural relations relevant…

Machine Learning · Computer Science 2026-03-03 Amir Asiaee , Zhuohui J. Liang , Chao Yan