English
Related papers

Related papers: Stop Using the Wilcoxon Test: Myth, Misconception …

200 papers

Attacks on the P-value are nothing new, but the recent attacks are increasingly more serious. They come from more mainstream sources, with widening targets such as a call to retire the significance testing altogether. While well meaning, I…

Other Statistics · Statistics 2022-01-11 Yudi Pawitan

One-sided t-tests are commonly used in the neuroimaging field, but two-sided tests should be the default unless a researcher has a strong reason for using a one-sided test. Here we extend our previous work on cluster false positive rates,…

Applications · Statistics 2019-01-03 Anders Eklund , Hans Knutsson , Thomas E. Nichols

Inverse normal transformations applied to the partially overlapping samples t-tests by Derrick et.al. (2017) are considered for their Type I error robustness and power. The inverse normal transformation solutions proposed in this paper are…

Computation · Statistics 2017-08-02 Ben Derrick , Paul White , Deirdre Toher

We analyze different types of simulations that applied researchers can use to assess whether their inference methods reliably control false-positive rates. We show that different assessments involve trade-offs, varying in the types of…

Econometrics · Economics 2025-10-03 Bruno Ferman

A crucial issue of current text generation models is that they often uncontrollably generate factually inconsistent text with respective of their inputs. Limited by the lack of annotated data, existing works in evaluating factual…

Computation and Language · Computer Science 2023-05-30 Wenhao Wu , Wei Li , Xinyan Xiao , Jiachen Liu , Sujian Li , Yajuan Lv

Nonparametric generalized likelihood ratio test is popularly used for model checking for regressions. However, there are two issues that may be the barriers for its powerfulness. First, the bias term in its liming null distribution causes…

Methodology · Statistics 2015-07-23 Cuizhen Niu , Xu Guo , Lixing Zhu

In this paper we proposed the alternative test to the two independent and normally distributed samples t test based on the cross variance concept. We present the simulation results of the power and the error rate of the special case of the…

Methodology · Statistics 2015-01-27 Rohmatul Fajriyah

Multiple hypothesis testing (MHT) frequently arises in scientific inquiries, and concurrent testing of multiple hypotheses inflates the risk of Type-I errors or false positives, rendering MHT corrections essential. This paper addresses MHT…

Ranks estimated from data are uncertain and this poses a challenge in many applications. However, estimated ranks are deterministic functions of estimated parameters, so the uncertainty in the ranks must be determined by the uncertainty in…

Methodology · Statistics 2023-06-22 Justin Rising

Web Applications (WA's) failures may lead to collapse of the institutions, therefore the importance of good quality WA's is increasing over the time. Testing is one of the best quality metrics that decide whether WA's are reliable or not.…

Software Engineering · Computer Science 2012-05-31 Mostafa Kandil , Ehab Hassanein , Sherif Mazen

Learning to Rank (LTR) methods are vital in online economies, affecting users and item providers. Fairness in LTR models is crucial to allocate exposure proportionally to item relevance. Widely used deterministic LTR models can lead to…

Machine Learning · Computer Science 2024-05-21 Ruocheng Guo , Jean-François Ton , Yang Liu , Hang Li

The one-sample log-rank test is the method of choice for single-arm Phase II trials with time-to-event endpoint. It allows to compare the survival of the patients to a reference survival curve that typically represents the expected survival…

Methodology · Statistics 2026-03-02 Jannik Feld , Moritz Fabian Danzer , Andreas Faldum , Rene Schmidt

Simulation has emerged as a popular method to study the long-term societal consequences of recommender systems. This approach allows researchers to specify their theoretical model explicitly and observe the evolution of system-level…

Computers and Society · Computer Science 2021-07-29 Eli Lucherini , Matthew Sun , Amy Winecoff , Arvind Narayanan

Code Review (CR) is the cornerstone for software quality assurance and a crucial practice for software development. As CR research matures, it can be difficult to keep track of the best practices and state-of-the-art in methodology,…

Software Engineering · Computer Science 2021-04-14 Dong Wang , Yuki Ueda , Raula Gaikovina Kula , Takashi Ishio , Kenichi Matsumoto

Score-based tests have been used to study parameter heterogeneity across many types of statistical models. This chapter describes a new self-normalization approach for score-based tests of mixed models, which addresses situations where…

Methodology · Statistics 2023-06-13 Ting Wang , Edgar Merkle

The Improbability Scale (IS) is proposed as a way of communicating to the general public the improbability (and by implication, the probability) of events predicted as the result of scientific research. Through the use of the Improbability…

Physics and Society · Physics 2007-05-23 David J. Ritchie

Bibliographic metrics are commonly utilized for evaluation purposes within academia, often in conjunction with other metrics. These metrics vary widely across fields and change with the seniority of the scholar; consequently, the only way…

Digital Libraries · Computer Science 2021-12-17 Sen Tian , Panos Ipeirotis

I present a critique of the methods used in a typical paper. This leads to three broad conclusions about the conventional use of statistical methods. First, results are often reported in an unnecessarily obscure manner. Second, the null…

Applications · Statistics 2013-03-05 Michael Wood

Rank correlations have found many innovative applications in the last decade. In particular, suitable rank correlations have been used for consistent tests of independence between pairs of random variables. Using ranks is especially…

Statistics Theory · Mathematics 2021-05-04 Hongjian Shi , Marc Hallin , Mathias Drton , Fang Han

Failures of retraction are common in science. Why do these failures occur? And, relatedly, what makes findings harder or easier to retract? We use data from Microsoft Academic Graph, Retraction Watch, and Altmetric -- including retracted…

Digital Libraries · Computer Science 2025-04-23 Shahan Ali Memon , Jevin D. West , Cailin O'Connor