English
Related papers

Related papers: Stop Using the Wilcoxon Test: Myth, Misconception …

200 papers

Hypothesis tests are a crucial statistical tool for data mining and are the workhorse of scientific research in many fields. Here we study differentially private tests of independence between a categorical and a continuous variable. We take…

Methodology · Statistics 2019-03-25 Simon Couch , Zeki Kazan , Kaiyan Shi , Andrew Bray , Adam Groce

Generative surveying -- where collections of LLM-based personas provide feedback on messages -- has emerged as a cheap and scalable alternative to traditional market research. However, LLMs are sensitive to small variations in prompt design…

Methodology · Statistics 2026-05-28 Hayden Helm , Carey Priebe

Much of science is (rightly or wrongly) driven by hypothesis testing. Even in situations where the hypothesis testing paradigm is correct, the common practice of basing inferences solely on p-values has been under intense criticism for over…

Methodology · Statistics 2015-12-31 M. J. Bayarri , Daniel J. Benjamin , James O. Berger , Thomas M. Sellke

Traditional statistics forbids use of test data (a.k.a. holdout data) during training. Dwork et al. 2015 pointed out that current practices in machine learning, whereby researchers build upon each other's models, copying hyperparameters and…

Machine Learning · Computer Science 2021-03-01 Sanjeev Arora , Yi Zhang

This research investigates how to determine whether two rankings come from the same distribution. We evaluate three hybrid tests: Wilcoxon's, Dietterich's, and Alpaydin's statistical tests combined with cross-validation (CV), each operating…

Methodology · Statistics 2022-02-14 Balázs R. Sziklai , Máté Baranyi , Károly Héberger

In this paper, we propose a power comparison between high dimensional t-test, sign and signed rank test for the one sample mean test. We show that the high dimensional signed rank test is superior to a high dimensional t test, but inferior…

Methodology · Statistics 2018-12-31 Long Feng

A fundamental challenge in comparing two survival distributions with right censored data is the selection of an appropriate nonparametric test, as the power of standard tests like the Log rank and Wilcoxon is highly dependent on the often…

Methodology · Statistics 2025-10-09 Abid Hussain , Touqeer Ahmad

Given independent samples from two univariate distributions, the one-sided Wilcoxon-Mann-Whitney statistic may be used to conduct a rank-based test of first-order stochastic dominance. We broaden the scope of applicability of such tests by…

Econometrics · Economics 2026-03-03 Brendan K. Beare , Jackson D. Clarke

Statistical hypothesis testing is the central method to demarcate scientific theories in both exploratory and inferential analyses. However, whether this method befits such purpose remains a matter of debate. Established approaches to…

Data Analysis, Statistics and Probability · Physics 2024-10-29 Orestis Loukas , Ho-Ryun Chung

Following discussions in 2010 and 2011, scientometric evaluators have increasingly abandoned relative indicators in favor of comparing observed with expected citation ratios. The latter method provides parameters with error values allowing…

Digital Libraries · Computer Science 2018-08-30 Loet Leydesdorff , Tobias Opthof

The current state of evaluation in survival analysis is plagued by the persistent use of evaluation metrics in ways that are misaligned with the stated modeling objective. In addition, many such evaluations are based on censoring…

Properly benchmarking a system is a difficult and intricate task. Unfortunately, even a seemingly innocuous benchmarking mistake can compromise the guarantees provided by a given systems security defense and also put its reproducibility and…

Cryptography and Security · Computer Science 2018-01-09 Erik van der Kouwe , Dennis Andriesse , Herbert Bos , Cristiano Giuffrida , Gernot Heiser

The inflation of Type I error rates is thought to be one of the causes of the replication crisis. Questionable research practices such as p-hacking are thought to inflate Type I error rates above their nominal level, leading to unexpectedly…

Methodology · Statistics 2024-12-31 Mark Rubin

Detecting and locating changes in highly multivariate data is a major concern in several current statistical applications. In this context, the first contribution of the paper is a novel non-parametric two-sample homogeneity test for…

Statistics Theory · Mathematics 2012-02-13 Alexandre Lung-Yut-Fong , Céline Lévy-Leduc , Olivier Cappé

The introduction of checkpoint inhibitors in immuno-oncology has raised questions about the suitability of the log-rank test as the default primary analysis method in confirmatory studies, particularly when survival curves exhibit…

Methodology · Statistics 2024-12-20 Dominic Magirr , Fredrik Öhrn

The F-measure or F-score is one of the most commonly used single number measures in Information Retrieval, Natural Language Processing and Machine Learning, but it is based on a mistake, and the flawed assumptions render it unsuitable for…

Information Retrieval · Computer Science 2019-09-13 David M. W. Powers

Like it or not, attempts to evaluate and monitor the quality of academic research have become increasingly prevalent worldwide. Performance reviews range from at the level of individuals, through research groups and departments, to entire…

Physics and Society · Physics 2017-03-31 R. Kenna , O. Mryglod , B. Berche

Users rely on search engines for information in critical contexts, such as public health emergencies. Understanding how users evaluate the trustworthiness of search results is therefore essential. Research has identified rank and the…

Human-Computer Interaction · Computer Science 2023-09-21 Sterling Williams-Ceci , Michael Macy , Mor Naaman

This paper calls attention to the missing component of the recommender system evaluation process: Statistical Inference. There is active research in several components of the recommender system evaluation process: selecting baselines,…

Information Retrieval · Computer Science 2021-09-15 Ngozi Ihemelandu , Michael D. Ekstrand

Given the well-known and fundamental problems with hypothesis testing via classical (point-form) significance tests, there has been a general move to alternative approaches, often focused on the Bayesian t-test. We show that the Bayesian…

Statistics Theory · Mathematics 2022-11-07 Fintan Costello , Paul Watts