English
Related papers

Related papers: Stop Using the Wilcoxon Test: Myth, Misconception …

200 papers

Moffat recently commented on our previous work. Our work focused on how laying the foundations of our evaluation methodology into the theory of measurement can improve our knowledge and understanding of the evaluation measures we use in IR…

Information Retrieval · Computer Science 2022-12-23 Marco Ferrante , Nicola Ferro , Norbert Fuhr

Alerting experience with a well-acknowledged safety analysis code initiated the authors to pay attention to safety issues of complex systems. Their first concern was the statistical characteristics of such a code. We point out a remarkable…

Data Analysis, Statistics and Probability · Physics 2007-05-23 L. Pal , M. Makai

In general, the more measurements we perform, the more information we gain about the system and thus, the more adequate decisions we will be able to make. However, in situations when we perform measurements to check for safety, the…

Other Statistics · Statistics 2023-05-16 Hector Reyes , Saeid Tizpaz-Niari , Vladik Kreinovich

Most researchers acknowledge an intrinsic hierarchy in the scholarly journals ('journal rank') that they submit their work to, and adjust not only their submission but also their reading strategies accordingly. On the other hand, much has…

Digital Libraries · Computer Science 2013-05-13 Björn Brembs , Marcus Munafò

The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluation of toxicity benchmarks. As organizations increasingly rely on these benchmarks to…

Artificial Intelligence · Computer Science 2026-05-12 Regina Gugg , Selina Niederländer , Andreas Stöckl , Martin Flechl

The Implicit Association Test, IAT, is widely used to measure hidden (subconscious) human biases, implicit bias, of many topics: race, gender, age, ethnicity, religion stereotypes. There is a need to understand the reliability of these…

Applications · Statistics 2023-12-27 S. Stanley Young , Warren B. Kindzierski

This paper considers the problem of directly generalizing the R-estimator under a linear model formulation with right-censored outcomes. We propose a natural generalization of the rank and corresponding estimating equation for the…

Methodology · Statistics 2026-01-13 Glen A. Satten , Mo Li , Ni Zhao , Robert L. Strawderman

Cross-validation is a popular non-parametric method for evaluating the accuracy of a predictive rule. The usefulness of cross-validation depends on the task we want to employ it for. In this note, I discuss a simple non-parametric setting,…

Methodology · Statistics 2019-09-27 Stefan Wager

Rerankers, typically cross-encoders, are computationally intensive but are frequently used because they are widely assumed to outperform cheaper initial IR systems. We challenge this assumption by measuring reranker performance for full…

Information Retrieval · Computer Science 2025-07-14 Mathew Jacob , Erik Lindgren , Matei Zaharia , Michael Carbin , Omar Khattab , Andrew Drozdov

Comparative simulation studies are workhorse tools for benchmarking statistical methods. As with other empirical studies, the success of simulation studies hinges on the quality of their design, execution and reporting. If not conducted…

Methodology · Statistics 2023-03-10 Samuel Pawel , Lucas Kook , Kelly Reeve

Constructed-response (CR) questions are a mainstay of introductory physics textbooks and exams. However, because of time, cost, and scoring reliability constraints associated with this format, CR questions are being increasingly replaced by…

Physics Education · Physics 2015-06-22 Aaron D. Slepkov , Ralph C. Shiell

Proportional hazards are a common assumption when designing confirmatory clinical trials in oncology. With the emergence of immunotherapy and novel targeted therapies, departure from the proportional hazard assumption is not rare in…

Methodology · Statistics 2020-08-27 José L. Jiménez

To evaluate Information Retrieval (IR) effectiveness, a possible approach is to use test collections, which are composed of a collection of documents, a set of description of information needs (called topics), and a set of relevant…

Information Retrieval · Computer Science 2020-11-03 Kevin Roitero

In this paper we argue that conventional unitary-invariant measures of recommender system (RS) performance based on measuring differences between predicted ratings and actual user ratings fail to assess fundamental RS properties. More…

Information Retrieval · Computer Science 2024-04-29 Tung Nguyen , Jeffrey Uhlmann

Predictive models are often required to produce reliable predictions under statistical conditions that are not matched to the training data. A common type of training-testing mismatch is covariate shift, where the conditional distribution…

Machine Learning · Computer Science 2025-01-22 Matteo Zecchin , Fredrik Hellström , Sangwoo Park , Shlomo Shamai , Osvaldo Simeone

Much research on Machine Learning testing relies on empirical studies that evaluate and show their potential. However, in this context empirical results are sensitive to a number of parameters that can adversely impact the results of the…

Software Engineering · Computer Science 2023-09-12 Salah Ghamizi , Maxime Cordy , Yuejun Guo , Mike Papadakis , And Yves Le Traon

While reliable data-driven decision-making hinges on high-quality labeled data, the acquisition of quality labels often involves laborious human annotations or slow and expensive scientific measurements. Machine learning is becoming an…

Machine Learning · Statistics 2024-03-01 Tijana Zrnic , Emmanuel J. Candès

An implicit association test is a human psychological test used to measure subconscious associations. While widely recognized by psychologists as an effective tool in measuring attitudes and biases, the validity of the results can be…

Human-Computer Interaction · Computer Science 2019-09-04 Brendon Boldt , Zack While , Eric Breimer

Randomized benchmarking (RB) is a popular procedure used to gauge the performance of a set of gates useful for quantum information processing (QIP). Recently, Proctor et al. [Phys. Rev. Lett. 119, 130502 (2017)] demonstrated a practically…

Quantum Physics · Physics 2019-07-04 Jiaan Qi , Hui Khoon Ng

Generalized linear models are often misspecified due to overdispersion, heteroscedasticity and ignored nuisance variables. Existing quasi-likelihood methods for testing in misspecified models often do not provide satisfactory type-I error…

Methodology · Statistics 2020-05-13 Jesse Hemerik , Jelle J Goeman , Livio Finos
‹ Prev 1 3 4 5 6 7 10 Next ›