Related papers: Statistical Tests and Research Assessments: A comm…
Measuring science is based on comparing articles to similar others. However, keyword-based groups of thematically similar articles are dominantly small. These small sizes keep the statistical errors of comparisons high. With the growing…
Hierarchically-organized data arise naturally in many psychology and neuroscience studies. As the standard assumption of independent and identically distributed samples does not hold for such data, two important problems are to accurately…
Alternative metrics (aka altmetrics) are gaining increasing interest in the scientometrics community as they can capture both the volume and quality of attention that a research work receives online. Nevertheless, there is limited knowledge…
What proportion of treated units actually benefited from an experimental intervention? What is the median or the largest individual treatment effect? This paper develops methods for answering such questions about the distribution of…
The size-dependent nature of the so-called group or departmental h-index is reconsidered in this paper. While the influence of unit size on such collective measures was already demonstrated a decade ago, institutional ratings based on this…
Novel reinforcement learning algorithms, or improvements on existing ones, are commonly justified by evaluating their performance on benchmark environments and are compared to an ever-changing set of standard algorithms. However, despite…
The launch of Google Scholar Metrics as a tool for assessing scientific journals may be serious competition for Thomson Reuters Journal Citation Reports, and for Scopus powered Scimago Journal Rank. A review of these bibliometric journal…
Given the growing use of impact metrics in the evaluation of scholars, journals, academic institutions, and even countries, there is a critical need for means to compare scientific impact across disciplinary boundaries. Unfortunately,…
Properties of a percentile-based rating scale needed in bibliometrics are formulated. Based on these properties, P100 was recently introduced as a new citation-rank approach (Bornmann, Leydesdorff, & Wang, in press). In this paper, we…
We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real research software ecosystems. Unlike existing scientific…
Assessing the research performance of multi-disciplinary institutions, where scientists belong to many fields, requires that the evaluators plan how to aggregate the performance measures of the various fields. Two methods of aggregation are…
With limited resources, scientific inquiries must be prioritized for further study, funding, and translation based on their practical significance: whether the effect size is large enough to be meaningful in the real world. Doing so must…
I present a critique of the methods used in a typical paper. This leads to three broad conclusions about the conventional use of statistical methods. First, results are often reported in an unnecessarily obscure manner. Second, the null…
Impact factors (and similar measures such as the Scimago Journal Rankings) suffer from two problems: (i) citation behavior varies among fields of science and therefore leads to systematic differences, and (ii) there are no statistics to…
Once upon a time, scientists' worth was measured by their ideas, proofs, and perhaps how eloquently they debated Hilbert's problems at seminars. But now, citation metrics have come to center stage and handed us new masters: FWCI and CNCI.…
In order to advance academic research, it is important to assess and evaluate the academic influence of researchers and the findings they produce. Citation metrics are universally used methods to evaluate researchers. Amongst the several…
Scholarly usage data holds the potential to be used as a tool to study the dynamics of scholarship in real time, and to form the basis for the definition of novel metrics of scholarly impact. However, the formal groundwork to reliably and…
We examine the role of trustworthiness and trust in statistical inference, arguing that it is the extent of trustworthiness in inferential statistical tools which enables trust in the conclusions. Certain tools, such as the p-value and…
To provide users insight into the value and limits of world university rankings, a comparative analysis is conducted of 5 ranking systems: ARWU, Leiden, THE, QS and U-Multirank. It links these systems with one another at the level of…
In many biological applications, the primary objective of study is to quantify the magnitude of treatment effect between two groups. Cohens'd or strictly standardized mean difference (SSMD) can be used to measure effect size however, it is…