English
Related papers

Related papers: Estimating the standard error of cross-Validation-…

200 papers

We propose a novel sampling-based federated learning framework for statistical inference on M-estimators with non-smooth objective functions, which frequently arise in modern statistical applications such as quantile regression and AUC…

Methodology · Statistics 2025-05-06 Xiudi Li , Lu Tian , Tianxi Cai

Modern causal inference methods allow machine learning to be used to weaken parametric modeling assumptions. However, the use of machine learning may result in complications for inference. Doubly-robust cross-fit estimators have been…

Methodology · Statistics 2022-03-11 Paul N Zivich , Alexander Breskin

Statistical inference methods are fundamentally important in machine learning. Most state-of-the-art inference algorithms are variants of Markov chain Monte Carlo (MCMC) or variational inference (VI). However, both methods struggle with…

Machine Learning · Computer Science 2019-10-17 Yichuan Zhang , José Miguel Hernández-Lobato

When performing supervised learning with the model selected using validation error from sample splitting and cross validation, the minimum value of the validation error can be biased downward. We propose two simple methods that use the…

Methodology · Statistics 2018-02-13 Leying Guan

The recent growth in multi-fidelity uncertainty quantification has given rise to a large set of variance reduction techniques that leverage information from model ensembles to provide variance reduction for estimates of the statistics of a…

Methodology · Statistics 2021-01-11 Trung Pham , Alex A. Gorodetsky

Bootstrap smoothed (bagged) parameter estimators have been proposed as an improvement on estimators found after preliminary data-based model selection. The key result of Efron (2014) is a very convenient and widely applicable formula for a…

Methodology · Statistics 2019-04-29 Paul Kabaila , Christeen Wijethunga

When evaluating and comparing models using leave-one-out cross-validation (LOO-CV), the uncertainty of the estimate is typically assessed using the variance of the sampling distribution. Considering the uncertainty is important, as the…

Methodology · Statistics 2022-02-16 Tuomas Sivula , Måns Magnusson , Aki Vehtari

Various risk-limiting audit (RLA) methods have been developed for instant-runoff voting (IRV) elections. A recent method, AWAIRE, is the first efficient approach that can take advantage of but does not require cast vote records (CVRs).…

Computers and Society · Computer Science 2024-12-03 Alexander Ek , Philip B. Stark , Peter J. Stuckey , Damjan Vukcevic

We propose a difference-based nonparametric methodology for the estimation and inference of the time-varying auto-covariance functions of a locally stationary time series when it is contaminated by a complex trend with both abrupt and…

Statistics Theory · Mathematics 2020-03-12 Yan Cui , Michael Levine , Zhou Zhou

Robins et al. (2008, 2017) applied the theory of higher order influence functions (HOIFs) to derive an estimator of the mean $\psi$ of an outcome Y in a missing data model with Y missing at random conditional on a vector X of continuous…

Statistics Theory · Mathematics 2026-01-27 Lin Liu , Rajarshi Mukherjee , Whitney K. Newey , James M. Robins

In recent years, data selection has emerged as a core issue for large-scale visual-language model pretraining, especially on noisy web-curated datasets. One widely adopted strategy assigns quality scores such as CLIP similarity for each…

Machine Learning · Computer Science 2024-02-06 Yiping Wang , Yifang Chen , Wendan Yan , Kevin Jamieson , Simon Shaolei Du

Multivariable predictive models are important statistical tools for providing synthetic diagnosis and prognostic algorithms based on multiple patients' characteristics. Their apparent discriminant and calibration measures usually have…

Applications · Statistics 2021-07-14 Katsuhiro Iba , Tomohiro Shinozaki , Kazushi Maruo , Hisashi Noma

For discrete-valued time series, predictive inference cannot be implemented through the construction of prediction intervals to some predetermined coverage level, as this is the case for real-valued time series. To address this problem, we…

Methodology · Statistics 2025-07-23 Maxime Faymonville , Carsten Jentsch , Efstathios Paparoditis

We present a practical approach for computing the sandwich variance estimator in two-stage regression model settings. As a motivating example for two-stage regression, we consider regression calibration, a popular approach for addressing…

Methodology · Statistics 2022-09-22 Lillian A. Boe , Thomas Lumley , Pamela A. Shaw

We develop tools to do valid post-selective inference for a family of model selection procedures, including choosing a model via cross-validated Lasso. The tools apply universally when the following random vectors are jointly asymptotically…

Methodology · Statistics 2018-02-13 Jelena Markovic , Lucy Xia , Jonathan Taylor

This paper proposes a closed-form optimal estimator based on the theory of estimating functions for a class of linear ARCH models. The estimating function (EF) estimator has the advantage over the widely used maximum likelihood (ML) and…

Statistics Theory · Mathematics 2008-12-05 Ajay Chandra

We introduce a generic estimator for the false discovery rate of any model selection procedure, in common statistical modeling settings including the Gaussian linear model, Gaussian graphical model, and model-X setting. We prove that our…

Methodology · Statistics 2026-02-25 Yixiang Luo , William Fithian , Lihua Lei

Instrumental variable models allow us to identify a causal function between covariates $X$ and a response $Y$, even in the presence of unobserved confounding. Most of the existing estimators assume that the error term in the response $Y$…

Machine Learning · Statistics 2022-09-23 Sorawit Saengkyongam , Leonard Henckel , Niklas Pfister , Jonas Peters

In machine learning one often assumes the data are independent when evaluating model performance. However, this rarely holds in practise. Geographic information data sets are an example where the data points have stronger dependencies among…

Applications · Statistics 2020-06-01 Jonne Pohjankukka , Tapio Pahikkala , Paavo Nevalainen , Jukka Heikkonen

Stratification is commonly employed in clinical trials to reduce the chance covariate imbalances and increase the precision of the treatment effect estimate. We propose a general framework for constructing the confidence interval (CI) for a…

Methodology · Statistics 2021-10-26 Yongqiang Tang
‹ Prev 1 8 9 10 Next ›