English
Related papers

Related papers: The Accuracy of Confidence Intervals for Field Nor…

200 papers

Purpose: Analyze the diversity of citation distributions to publications in different research topics to investigate the accuracy of size-independent, rank-based indicators. Top percentile-based indicators are the most common indicators of…

Digital Libraries · Computer Science 2024-07-15 Alonso Rodríguez-Navarro

Estimating the prevalence of a category in a population using imperfect measurement devices (diagnostic tests, classifiers, or large language models) is fundamental to science, public health, and online trust and safety. Standard approaches…

Artificial Intelligence · Computer Science 2026-04-24 Fridolin Linder , Thomas Leeper , Daniel Haimovich , Niek Tax , Lorenzo Perini , Milan Vojnovic

The two most used citation impact indicators in the assessment of scientific journals are, nowadays, the impact factor and the h-index. However, both indicators are not field normalized (vary heavily depending on the scientific category)…

Digital Libraries · Computer Science 2015-10-14 Sara M. Gonzalez-Betancor , Pablo Dorta-Gonzalez

Measuring the impact of a publication in a fair way is a significant challenge in bibliometrics, as it must not introduce biases between fields and should enable comparison of the impact of publications from different years. In this paper,…

Digital Libraries · Computer Science 2024-03-07 Emilio Gómez-Déniz , Pablo Dorta-González

This paper proposes a new non-parametric bootstrap method to quantify the uncertainty of average treatment effect estimate for the treated from matching estimators. More specifically, it seeks to quantify the uncertainty associated with the…

Methodology · Statistics 2024-08-21 Jing Li

Random-effects meta-analyses have been widely applied in evidence synthesis for various types of medical studies. However, standard inference methods (e.g. restricted maximum likelihood estimation) usually underestimate statistical errors…

Methodology · Statistics 2019-05-13 Shonosuke Sugasawa , Hisashi Noma

Knowing when a classifier's prediction can be trusted is useful in many applications and critical for safely using AI. While the bulk of the effort in machine learning research has been towards improving classifier performance,…

Machine Learning · Statistics 2018-10-30 Heinrich Jiang , Been Kim , Melody Y. Guan , Maya Gupta

When analyzing incomplete data, is it better to use multiple imputation (MI) or full information maximum likelihood (ML)? In large samples ML is clearly better, but in small samples ML's usefulness has been limited because ML commonly uses…

Methodology · Statistics 2017-03-24 Paul T. von Hippel

This paper argues for the widest possible use of bootstrap confidence intervals for comparing NLP system performances instead of the state-of-the-art status (SOTA) and statistical significance testing. Their main benefits are to draw…

Computation and Language · Computer Science 2022-05-24 Yves Bestgen

For differences between means of continuous data from independent groups, the customary scale-free measure of effect is the standardized mean difference (SMD). To justify use of SMD, one should be reasonably confident that the group-level…

Statistics Theory · Mathematics 2025-12-10 Elena Kulinskaya , David C. Hoaglin

The problem of quantifying uncertainty about the locations of multiple change points by means of confidence intervals is addressed. The asymptotic distribution of the change point estimators obtained as the local maximisers of moving sum…

Methodology · Statistics 2022-06-20 Haeran Cho , Claudia Kirch

The quality of a summarization evaluation metric is quantified by calculating the correlation between its scores and human annotations across a large number of summaries. Currently, it is unclear how precise these correlation estimates are,…

Computation and Language · Computer Science 2021-07-28 Daniel Deutsch , Rotem Dror , Dan Roth

Linear combinations of multinomial probabilities, such as those resulting from contingency tables, are of use when evaluating classification system performance. While large sample inference methods for these combinations exist, small sample…

Methodology · Statistics 2021-04-20 Katherine A. Batterton , Christine M. Schubert , Richard L. Warr

The bootstrap is a popular and convenient method for quantifying the authority of an empirical ordering of attributes, for example of a ranking of the performance of institutions or of the influence of genes on a response variable. In the…

Statistics Theory · Mathematics 2009-11-20 Peter Hall , Hugh Miller

This paper explores a new indicator of journal citation impact, denoted as source normalized impact per paper (SNIP). It measures a journal's contextual citation impact, taking into account characteristics of its properly defined subject…

Digital Libraries · Computer Science 2009-11-16 Henk F. Moed

Most NLP datasets are not annotated with protected attributes such as gender, making it difficult to measure classification bias using standard measures of fairness (e.g., equal opportunity). However, manually annotating a large dataset…

Computation and Language · Computer Science 2020-04-28 Kawin Ethayarajh

Citation numbers are extensively used for assessing the quality of scientific research. The use of raw citation counts is generally misleading, especially when applied to cross-disciplinary comparisons, since the average number of citations…

Physics and Society · Physics 2011-11-28 Filippo Radicchi , Claudio Castellano

A new method is proposed for the correction of confidence intervals when the original interval does not have the correct nominal coverage probabilities in the frequentist sense. The proposed method is general and does not require any…

Computation · Statistics 2013-08-30 P. Menendez , Y. Fan , P. H. Garthwaite , S. A. Sisson

Using the CD-ROM version of the Science Citation Index 2010 (N = 3,705 journals), we study the (combined) effects of (i) fractional counting on the impact factor (IF) and (ii) transformation of the skewed citation distributions into a…

Digital Libraries · Computer Science 2012-06-19 Loet Leydesdorff , Ping Zhou , Lutz Bornmann

Practical or scientific considerations often lead to selecting a subset of parameters as ``important.'' Inferences about those parameters often are based on the same data used to select them in the first place. That can make the reported…

Methodology · Statistics 2019-06-04 Yoav Benjamini , Yotam Hechtlinger , Philip B. Stark