English
Related papers

Related papers: The pigeonhole bootstrap

200 papers

Meta-analysis is a statistical method used in evidence synthesis for combining, analyzing and summarizing studies that have the same target endpoint and aims to derive a pooled quantitative estimate using fixed and random effects models or…

Methodology · Statistics 2022-04-25 Ivette Raices Cruz , Matthias C. M. Troffaes , Johan Lindström , Ullrika Sahlin

To draw scientifically meaningful conclusions and build reliable models of quantitative phenomena, cause and effect must be taken into consideration (either implicitly or explicitly). This is particularly challenging when the measurements…

Machine Learning · Computer Science 2020-12-11 Max A. Little , Reham Badawy

A general notion of bootstrapped $\phi$-divergence estimates constructed by exchangeably weighting sample is introduced. Asymptotic properties of these generalized bootstrapped $\phi$-divergence estimates are obtained, by mean of the…

Statistics Theory · Mathematics 2019-03-06 Salim Bouzebda , Mohamed Cherfi

Although the methods of bagging and random forests are some of the most widely used prediction methods, relatively little is known about their algorithmic convergence. In particular, there are not many theoretical guarantees for deciding…

Statistics Theory · Mathematics 2019-07-23 Miles E. Lopes

Bootstrap is a popular methodology for simulating input uncertainty. However, it can be computationally expensive when the number of samples is large. We propose a new approach called \textbf{Orthogonal Bootstrap} that reduces the number of…

Methodology · Statistics 2024-05-02 Kaizhao Liu , Jose Blanchet , Lexing Ying , Yiping Lu

Bipartite data is common in data engineering and brings unique challenges, particularly when it comes to clustering tasks that impose on strong structural assumptions. This work presents an unsupervised method for assessing similarity in…

Machine Learning · Computer Science 2017-02-17 Aaron Gerow , Mingyang Zhou , Stan Matwin , Feng Shi

We numerically analyze the random matrix ensembles of real-symmetric matrices with column/row constraints for many system conditions e.g. disorder type, matrix-size and basis-connectivity. The results reveal a rich behavior hidden beneath…

Statistical Mechanics · Physics 2015-10-28 Suchetana Sadhukhan , Pragya Shukla

Regression models with crossed random effect errors can be very expensive to compute. The cost of both generalized least squares and Gibbs sampling can easily grow as $N^{3/2}$ (or worse) for $N$ observations. Papaspiliopoulos et al. (2020)…

Methodology · Statistics 2021-03-22 Swarnadip Ghosh , Trevor Hastie , Art B. Owen

We propose a bootstrap testing framework for a general class of hypothesis tests, which allows resampling under the null hypothesis as well as other forms of bootstrapping. We identify combinations of resampling schemes and bootstrap…

Statistics Theory · Mathematics 2025-12-12 Alexis Derumigny , Miltiadis Galanis , Wieger Schipper , Aad van der Vaart

We consider the challenges that arise when fitting complex ecological models to 'large' data sets. In particular, we focus on random effect models which are commonly used to describe individual heterogeneity, often present in ecological…

Methodology · Statistics 2022-05-17 Ruth King , Blanca Sarzo , Víctor Elvira

We analyze the complexity of Gibbs samplers for inference in crossed random effect models used in modern analysis of variance. We demonstrate that for certain designs the plain vanilla Gibbs sampler is not scalable, in the sense that its…

Computation · Statistics 2018-03-28 Omiros Papaspiliopoulos , Gareth O. Roberts , Giacomo Zanella

We propose Posterior Bootstrap, a set of algorithms extending Weighted Likelihood Bootstrap, to properly incorporate prior information and address the problem of model misspecification in Bayesian inference. We consider two approaches to…

Methodology · Statistics 2021-04-19 Emilia Pompe

Meta-analyses require an effect-size estimate and its corresponding sampling variance from primary studies. In some cases, estimators for the sampling variance of a given effect size statistic may not exist, necessitating the derivation of…

Machine learning has demonstrated remarkable prediction accuracy over i.i.d data, but the accuracy often drops when tested with data from another distribution. In this paper, we aim to offer another view of this problem in a perspective…

Machine Learning · Computer Science 2022-06-20 Haohan Wang , Zeyi Huang , Hanlin Zhang , Yong Jae Lee , Eric Xing

Practical inference procedures for quantile regression models of panel data have been a pervasive concern in empirical work, and can be especially challenging when the panel is observed over many time periods and temporal dependence needs…

Econometrics · Economics 2025-07-25 Antonio F. Galvao , Carlos Lamarche , Thomas Parker

Bootstrapping was designed to randomly resample data from a fixed sample using Monte Carlo techniques. However, the original sample itself defines a discrete distribution. Convolutional methods are well suited for discrete distributions,…

Methodology · Statistics 2021-07-19 Jared M. Clark , Richard L. Warr

In order to test if an unknown matrix has a given rank (null hypothesis), we consider the family of statistics that are minimum squared distances between an estimator and the manifold of fixed-rank matrix. Under the null hypothesis, every…

Statistics Theory · Mathematics 2013-01-09 François Portier , Bernard Delyon

Cross-validation is a widely-used technique to estimate prediction error, but its behavior is complex and not fully understood. Ideally, one would like to think that cross-validation estimates the prediction error for the model at hand, fit…

Methodology · Statistics 2024-03-12 Stephen Bates , Trevor Hastie , Robert Tibshirani

We propose a bootstrap procedure for data that may exhibit clustering in two or more dimensions. We use insights from the theory of generalized U-statistics to analyze the large-sample properties of statistics that are sample averages from…

Methodology · Statistics 2017-12-06 Konrad Menzel

AI/ML methods are increasingly used in economics to generate binary variables (or labels) via classification algorithms. When these generated variables are included as covariates in regressions, even small misclassification errors can…

Econometrics · Economics 2026-04-28 Timothy Christensen , Silvia Goncalves , Benoit Perron