English
Related papers

Related papers: The Accuracy of Confidence Intervals for Field Nor…

200 papers

This work develops formal statistical inference procedures for machine learning ensemble methods. Ensemble methods based on bootstrapping, such as bagging and random forests, have improved the predictive accuracy of individual trees, but…

Machine Learning · Statistics 2015-09-11 Lucas Mentch , Giles Hooker

Out-of-bag error is commonly used as an estimate of generalisation error in ensemble-based learning models such as random forests. We present confidence intervals for this quantity using the delta-method-after-bootstrap and the…

Methodology · Statistics 2022-01-28 Samyak Rajanala , Stephen Bates , Trevor Hastie , Robert Tibshirani

In assessing prediction accuracy of multivariable prediction models, optimism corrections are essential for preventing biased results. However, in most published papers of clinical prediction models, the point estimates of the prediction…

This article explores combinations of weighted bootstraps, like the Bayesian bootstrap, with the bootstrap $t$ method for setting approximate confidence intervals for the mean of a random variable in small samples. For this problem the…

Statistics Theory · Mathematics 2025-08-21 Art B. Owen

I have three goals in this article: (1) To show the enormous potential of bootstrapping and permutation tests to help students understand statistical concepts including sampling distributions, standard errors, bias, confidence intervals,…

Other Statistics · Statistics 2014-11-20 Tim Hesterberg

The purpose of the present paper is to assess the efficacy of confidence intervals for Rosenthal's fail-safe number. Although Rosenthal's estimator is highly used by researchers, its statistical properties are largely unexplored. First of…

Methodology · Statistics 2015-09-07 Konstantinos C. Fragkos , Michail Tsagris , Christos C. Frangos

Bootstrap smoothed (bagged) parameter estimators have been proposed as an improvement on estimators found after preliminary data-based model selection. The key result of Efron (2014) is a very convenient and widely applicable formula for a…

Methodology · Statistics 2019-04-29 Paul Kabaila , Christeen Wijethunga

In the recent paper [5], a Bayesian approach for constructing confidence intervals in monotone regression problems is proposed, based on credible intervals. We view this method from a frequentist point of view, and show that it corresponds…

Statistics Theory · Mathematics 2023-08-01 Piet Groeneboom , Geurt Jongbloed

One of the most commonly used methods for forming confidence intervals for statistical inference is the empirical bootstrap, which is especially expedient when the limiting distribution of the estimator is unknown. However, despite its…

Statistics Theory · Mathematics 2020-11-24 Morgane Austern , Vasilis Syrgkanis

Summing or averaging nonlinearly field-normalized citation counts is a common but methodologically problematic practice, as it violates mathematical principles. The issue originates from the nonlinear transformation, which disrupts the…

Digital Libraries · Computer Science 2025-11-05 Limi Tang

With the passage of more time from the original date of publication, the measure of the impact of scientific works using subsequent citation counts becomes more accurate. However the measurement of individual and organizational research…

Digital Libraries · Computer Science 2018-11-06 Giovanni Abramo , Tindaro Cicero , Ciriaco Andrea D'Angelo

Confidence interval of mean is often used when quoting statistics. The same rigor is often missing when quoting percentiles and tolerance or percentile intervals. This article derives the expression for confidence in percentiles of a sample…

Methodology · Statistics 2024-03-01 Sanjay M. Joshi

For the quantification of QoE, subjects often provide individual rating scores on certain rating scales which are then aggregated into Mean Opinion Scores (MOS). From the observed sample data, the expected value is to be estimated. While…

Methodology · Statistics 2018-06-05 Tobias Hossfeld , Poul E. Heegaard , Martin Varela , Lea Skorin-Kapov

If we want to assess whether the paper in question has had a particularly high or low citation impact compared to other papers, the standard practice in bibliometrics is to normalize citations in respect of the subject category and…

Digital Libraries · Computer Science 2013-07-30 Lutz Bornmann , Werner Marx , Andreas Barth

The behaviors of various confidence/credible interval constructions are explored, particularly in the region of low statistics where methods diverge most. We highlight a number of challenges, such as the treatment of nuisance parameters,…

Data Analysis, Statistics and Probability · Physics 2015-02-04 Steven D. Biller , Scott M. Oser

Introductory texts on statistics typically only cover the classical "two sigma" confidence interval for the mean value and do not describe methods to obtain confidence intervals for other estimators. The present technical report fills this…

Methodology · Statistics 2018-07-11 Christoph Dalitz

Journal field classifications in Scopus are used for citation-based indicators and by authors choosing appropriate journals to submit to. Whilst prior research has found that Scopus categories are occasionally misleading, it is not known…

Digital Libraries · Computer Science 2023-07-31 Mike Thelwall , Stephen Pinfield

Large language models (LLMs) are increasingly used in social science as scalable measurement tools for converting unstructured text into variables that can enter standard empirical designs. Measurement validity demands more than high…

Artificial Intelligence · Computer Science 2026-05-13 Jinyuan Wang , Ningyuan Deng , Yi Yang

It is widely recognized that citation counts for papers from different fields cannot be directly compared because different scientific fields adopt different citation practices. Citation counts are also strongly biased by paper age since…

Physics and Society · Physics 2017-08-30 Giacomo Vaccario , Matus Medo , Nicolas Wider , Manuel Sebastian Mariani

Bootstrapping is often applied to get confidence limits for semiparametric inference of a target parameter in the presence of nuisance parameters. Bootstrapping with replacement can be computationally expensive and problematic when…