Related papers: The Accuracy of Confidence Intervals for Field Nor…
We consider the problem of interval estimation of the odds ratio. An asymptotic confidence interval is widely applied in medical research. Unfortunately that confidence interval has a poor coverage probability: it is significantly smaller…
In this paper, we show that citation counts work better than a random baseline (by a margin of 10%) in distinguishing excellent research, while Mendeley reader counts don't work better than the baseline. Specifically, we study the potential…
In Natural Language Processing (NLP), binary classification algorithms are often evaluated using the F1 score. Because the sample F1 score is an estimate of the population F1 score, it is not sufficient to report the sample F1 score without…
As large language models (LLMs) are increasingly used in high-stakes domains, accurately assessing their confidence is crucial. Humans typically express confidence through epistemic markers (e.g., "fairly confident") instead of numerical…
A prediction interval covers a future observation from a random process in repeated sampling, and is typically constructed by identifying a pivotal quantity that is also an ancillary statistic. Analogously, a tolerance interval covers a…
We propose a methodology for constructing confidence regions with partially identified models of general form. The region is obtained by inverting a test of internal consistency of the econometric structure. We develop a dilation bootstrap…
The Leiden Rankings can be used for grouping research universities by considering universities which are not statistically significantly different as homogeneous sets. The groups and intergroup relations can be analyzed and visualized using…
Citation averages, and Impact Factors (IFs) in particular, are sensitive to sample size. We apply the Central Limit Theorem (CLT) to IFs to understand their scale-dependent behavior. For a journal of $n$ randomly selected papers from a…
Clustered data arise naturally in many scientific and applied research settings where units are grouped within clusters. They are commonly analyzed using linear mixed models to account for within-cluster correlations. This article focuses…
Benchmarking models is a key factor for the rapid progress in machine learning (ML) research. Thus, further progress depends on improving benchmarking metrics. A standard metric to measure the behavioral alignment between ML models and…
Bootstrapping and other resampling methods are increasingly appearing in the textbooks and curricula of courses that introduce undergraduate students to statistical methods. In order to teach the bootstrap well, students and instructors…
The ISO 5725 series frames interlaboratory precision through repeatability, between-laboratory, and reproducibility variances, yet practical guidance on deploying bootstrap methods within this one-way random-effects setting remains limited.…
Fractional counting of citations can improve on ranking of multi-disciplinary research units (such as universities) by normalizing the differences among fields of science in terms of differences in citation behavior. Furthermore,…
The purpose of this study is to compare the changing behavior of two counting methods (whole counting and whole-normalized counting) and inflation rate at country level research productivity and impact. For this, publication data on…
Confidence calibration assumes a unique ground-truth label per input, yet this assumption fails wherever annotators genuinely disagree. Post-hoc calibrators fitted on majority-voted labels, the standard single-label targets used in…
In the analysis of survey data it is of interest to estimate and quantify uncertainty about means or totals for each of several non-overlapping subpopulations, or areas. When the sample size for a given area is small, standard confidence…
Citation counts are widely used as indicators of research quality to support or replace human peer review and for lists of top cited papers, researchers, and institutions. Nevertheless, the relationship between citations and research…
Journal Impact Factors (IFs) can be considered historically as the first attempt to normalize citation distributions by using averages over two years. However, it has been recognized that citation distributions vary among fields of science…
We investigate the calibration of large language models' (LLMs') confidence across diverse tasks. The results of our preregistered study show that the current crop of LLMs are, like people, too sure they are right: confidence exceeds…
The recent paper "Simple confidence intervals for MCMC without CLTs" by J.S. Rosenthal, showed the derivation of a simple MCMC confidence interval using only Chebyshev's inequality, not CLT. That result required certain assumptions about…