English
Related papers

Related papers: Stop Using the Wilcoxon Test: Myth, Misconception …

200 papers

Real-life data are often non-IID due to complex distributions and interactions, and the sensitivity to the distribution of samples can differ among learning models. Accordingly, a key question for any supervised or unsupervised model is…

Machine Learning · Computer Science 2023-10-03 Zhilin Zhao , Longbing Cao

A common method for deriving non-parametric tests is to reformulate a parametric test in terms of sample ranks. Despite being distribution free (even in finite samples), the resulting tests often display remarkable asymptotic power…

Statistics Theory · Mathematics 2022-08-10 Dan D. Erdmann-Pham , Jonathan Terhorst , Yun S. Song

In survey panel research, anonymity of the participants is of great importance, as it must be ensured to prevent negative effects of participation as well as to maintain trust that the sensitive data that respondents provide is handled with…

Methodology · Statistics 2014-04-08 Kees T. Mulder

While the asymptotic relative efficiency (ARE) of Wilcoxon rank-based tests for location and regression with respect to their parametric Student competitors can be arbitrarily large, Hodges and Lehmann (1961) have shown that the ARE of the…

Statistics Theory · Mathematics 2013-05-22 Marc Hallin , Thomas Verdebout , Yvik Swan

Despite advances in Automatic Speech Recognition (ASR), transcription errors persist and require manual correction. Confidence scores, which indicate the certainty of ASR results, could assist users in identifying and correcting errors.…

Human-Computer Interaction · Computer Science 2025-03-20 Korbinian Kuhn , Verena Kersken , Gottfried Zimmermann

Software testing has often to be done under severe pressure due to limited resources and a challenging time schedule facing the demand to assure the fulfillment of the software requirements. In addition, testing should unveil those software…

Software Engineering · Computer Science 2019-12-30 Michael Felderer , Ina Schieferdecker

Experimental research methods describe standards to safeguard scientific integrity and reputability. These methods have been extensively integrated into traditional scientific disciplines and studied in the philosophy of science. The field…

Cryptography and Security · Computer Science 2019-05-20 Carrie Gardner , Abby Waliga , David Thaw , Sarah Churchman

A test oracle determines whether a system behaves correctly for a given input. Automatic testing techniques rely on an automated test oracle to test the system without user interaction. Important families of automated test oracles include…

Software Engineering · Computer Science 2022-10-21 Manuel Rigger , Zhendong Su

Generative Artificial Intelligence (GenAI) is now widespread in education, yet the efficacy of GenAI systems remains constrained by the quality and interpretation of the labeled data used to train and evaluate them. Studies commonly report…

Computers and Society · Computer Science 2026-04-01 Danielle R. Thomas , Conrad Borchers , Kirk P. Vanacore , Kenneth R. Koedinger , René F. Kizilcec

Despite bilingual speakers frequently using mixed-language queries in web searches, Information Retrieval (IR) research on them remains scarce. To address this, we introduce MiLQ, Mixed-Language Query test set, the first public benchmark of…

Information Retrieval · Computer Science 2025-10-21 Jonghwi Kim , Deokhyung Kang , Seonjeong Hwang , Yunsu Kim , Jungseul Ok , Gary Lee

The usefulness evaluation model proposed by Cole et al. in 2009 [2] focuses on the evaluation of interactive IR systems by their support towards the user's overall goal, sub goals and tasks. This is a more human focus of the IR evaluation…

Information Retrieval · Computer Science 2018-09-10 Daniel Hienert , Peter Mutschke

Diagnostic accuracy studies assess sensitivity and specificity of a new index test in relation to an established comparator or the reference standard. The development and selection of the index test is usually assumed to be conducted prior…

Methodology · Statistics 2022-08-30 Max Westphal , Antonia Zapf

A large fraction of papers in the climate literature includes erroneous uses of significance tests. A Bayesian analysis is presented to highlight the meaning of significance tests and why typical misuse occurs. It is concluded that a…

Atmospheric and Oceanic Physics · Physics 2016-08-24 Maarten H. P. Ambaum

Non-parametric approaches to test for trends in time series make use of the Mann-Kendall statistic. Based on asymptotic arguments, these tests assume that its distribution follows a Gaussian distribution, even for autocorrelated time…

Applications · Statistics 2026-04-17 Tristan Gamot , Nils Thibeau--Sutre , Tom J. M. Van Dooren

Traditionally, Machine Translation (MT) Evaluation has been treated as a regression problem -- producing an absolute translation-quality score. This approach has two limitations: i) the scores lack interpretability, and human annotators…

Computation and Language · Computer Science 2024-01-31 Ibraheem Muhammad Moosa , Rui Zhang , Wenpeng Yin

Background. The reliability paradox describes the empirical observation that cognitive tasks producing robust group-level effects often yield poor between-individual reliability. Existing approaches rely predominantly on the intraclass…

Methodology · Statistics 2026-05-26 Maria Westrin

Adopting the scientific method a theoretical model is proposed as foundation for information science and technology, extending the existing theory of signaling: a fact f becomes known in a physical system only following the success of a…

Other Computer Science · Computer Science 2009-02-01 A. P. Young

The recently proposed Euclidean index offers a novel approach to measure the citation impact of academic authors, in particular as an alternative to the h-index. We test if the index provides new, robust information, not covered by existing…

Digital Libraries · Computer Science 2017-03-08 Jens Peter Andersen

STATCHECK is an R algorithm designed to scan papers automatically for inconsistencies between test statistics and their associated p values (Nuijten et al., 2016). The goal of this comment is to point out an important and well-documented…

Quantitative Methods · Quantitative Biology 2017-11-27 Thomas Schmidt

Randomized benchmarking and variants thereof, which we collectively call RB+, are widely used to characterize the performance of quantum computers because they are simple, scalable, and robust to state-preparation and measurement errors.…

Quantum Physics · Physics 2019-06-05 Robin Harper , Ian Hincks , Chris Ferrie , Steven T. Flammia , Joel J. Wallman
‹ Prev 1 8 9 10 Next ›