Related papers: Caveats for using statistical significance tests i…
Consistently checking the statistical significance of experimental results is the first mandatory step towards reproducible science. This paper presents a hitchhiker's guide to rigorous comparisons of reinforcement learning algorithms.…
Scholarly usage data provides unique opportunities to address the known shortcomings of citation analysis. However, the collection, processing and analysis of usage data remains an area of active research. This article provides a review of…
Sensitivity analysis (SA) has much to offer for a very large class of applications, such as model selection, calibration, optimization, quality assurance and many others. Sensitivity analysis offers crucial contextual information regarding…
Should students be used as experimental subjects in software engineering? Given that students are in many cases readily available and cheap it is no surprise that the vast majority of controlled experiments in software engineering use them.…
Testing for normality is a widely used procedure in statistics and data analysis, often applied prior to employing methods that rely on the assumption of normally distributed data. While several existing tests target distributional…
Once upon a time, scientists' worth was measured by their ideas, proofs, and perhaps how eloquently they debated Hilbert's problems at seminars. But now, citation metrics have come to center stage and handed us new masters: FWCI and CNCI.…
Several questions of scientometrics parameters organization are considered. Two new indices for scientific works citation analysis are proposed. They provide more detailed and reliable scientific significance assessment of individual…
Propensity score plays a central role in causal inference, but its use is not limited to causal comparisons. As a covariate balancing tool, propensity score can be used for controlled descriptive comparisons between groups whose memberships…
This book critically analyses the value of citation data, altmetrics, and artificial intelligence to support the research evaluation of articles, scholars, departments, universities, countries, and funders. It introduces and discusses…
Citation analysis is widely used in research evaluation to assess the impact of scientific papers. These analyses rest on the assumption that citation decisions by authors are accurate, representing the flow of knowledge from cited to…
Test-negative designs are widely used for post-market evaluation of vaccine effectiveness, particularly in cases when randomized trials are not feasible. Differing from classical test-negative designs where only healthcare-seekers with…
This paper provides a statistical method to test whether a system that performs a binary sequential hypothesis test is optimal in the sense of minimizing the average decision times while taking decisions with given reliabilities. The…
Bibliometrics plays an increasingly important role in research evaluation. However, no gold standard exists for a set of reliable and valid (field-normalized) impact indicators in research evaluation. This opinion paper recommends that…
In a critical and provocative paper, Abramo and D'Angelo claim that commonly used scientometric indicators such as the mean normalized citation score (MNCS) are completely inappropriate as indicators of scientific performance. Abramo and…
Creating test collections for offline retrieval evaluation requires human effort to judge documents' relevance. This expensive activity motivated much work in developing methods for constructing benchmarks with fewer assessment costs. In…
Research instruments play significant roles in the construction of scientific knowledge, even though we have only acquired very limited knowledge about their lifecycles from quantitative studies. This paper aims to address this gap by…
Despite their importance in supporting experimental conclusions, standard statistical tests are often inadequate for research areas, like the life sciences, where the typical sample size is small and the test assumptions difficult to…
In this article, we consider the problem of simultaneous testing of hypotheses when the individual test statistics are not necessarily independent. Specifically, we consider the problem of simultaneous testing of point null hypotheses…
Many U.S. colleges now use test-optional admissions. A frequent claim is that by not seeing standardized test scores, a college can admit a student body it prefers, say with more diversity. But how can observing less information improve…
Intercurrent (post-treatment) events occur frequently in randomized trials, and investigators often express interest in treatment effects that suitably take account of these events. A naive conditioning on intercurrent events does not have…