English
Related papers

Related papers: Conditional Equivalence Testing: an alternative re…

200 papers

Statistical hypothesis testing serves as statistical evidence for scientific innovation. However, if the reported results are intentionally biased, hypothesis testing no longer controls the rate of false discovery. In particular, we study…

Methodology · Statistics 2018-10-12 Junpei Komiyama , Takanori Maehara

Many tasks revolve around editing a document, whether code or text. We formulate the revision similarity problem to unify a wide range of machine learning evaluation problems whose goal is to assess a revision to an existing document. We…

Machine Learning · Computer Science 2025-12-16 Nikolai Gruzinov , Ksenia Sycheva , Earl T. Barr , Alex Bezzubov

Computerized adaptive testing (CAT) is an interesting and promising approach to testing human abilities. In our research we use Bayesian networks to create a model of tested humans. We collected data from paper tests performed with grammar…

Artificial Intelligence · Computer Science 2017-03-28 Martin Plajner , Jiří Vomlel

A/B tests, also known as randomized controlled experiments (RCTs), are the gold standard for evaluating the impact of new policies, products, or decisions. However, these tests can be costly in terms of time and resources, potentially…

Machine Learning · Statistics 2025-01-03 Shima Nassiri , Mohsen Bayati , Joe Cooprider

We study the model selection problem in conditional average treatment effect (CATE) prediction. Unlike previous works on this topic, we focus on preserving the rank order of the performance of candidate CATE predictors to enable accurate…

Machine Learning · Statistics 2020-07-20 Yuta Saito , Shota Yasui

We investigate a confirmatory two stage adaptive procedure for assessing average bioequivalence and provide some insights to its theoretical properties. Effectively, we perform Two One-Sided Tests (TOST) to reach overall decision about each…

Methodology · Statistics 2022-03-18 Marie Louise Østerdal , Kyle Raymond , Christian Pipper

We introduce equivalence testing procedures for linear regression analyses. Such tests can be very useful for confirming the lack of a meaningful association between a continuous outcome and a continuous or binary predictor. Specifically,…

Methodology · Statistics 2023-05-17 Harlan Campbell

Meta-analysis is a powerful tool for assessing drug safety by combining treatment-related toxicological findings across multiple studies, as clinical trials are typically underpowered for detecting adverse drug effects. However, incomplete…

Background: The standard regulatory approach to assess replication success is the two-trials rule, requiring both the original and the replication study to be significant with effect estimates in the same direction. The sceptical p-value…

Methodology · Statistics 2025-05-01 Jeanette Köppe , Charlotte Micheloud , Stella Erdmann , Rachel Heyard , Leonhard Held

Randomization tests allow simple and unambiguous tests of null hypotheses, by comparing observed data to a null ensemble in which experimentally-controlled variables are randomly resampled. In behavioral and neuroscience experiments,…

Methodology · Statistics 2023-11-08 Kenneth D. Harris , Kevin J. Miller

Replication studies are increasingly conducted but there is no established statistical criterion for replication success. We propose a novel approach combining reverse-Bayes analysis with Bayesian hypothesis testing: a sceptical prior is…

Methodology · Statistics 2022-12-15 Samuel Pawel , Leonhard Held

Equivalence testing plays a key role in several domains, such as the development of generic medical products, which are therapeutically equivalent to brand-name drugs but with reduced cost and increased accessibility. Promoting access to…

Methodology · Statistics 2025-07-31 Luca Insolia , Yanyuan Ma , Younes Boulaguiem , Stéphane Guerrier

Effective decision making from randomised controlled clinical trials relies on robust interpretation of the numerical results. However, the language we use to describe clinical trials can cause confusion both in trial design and in…

Computerized Adaptive Testing (CAT) offers an efficient and personalized method for assessing examinee proficiency by dynamically adjusting test questions based on individual performance. Compared to traditional, non-personalized testing…

Testing the equivalence of multiple quantiles between two populations is important in many scientific applications, such as clinical trials, where conventional mean-based methods may be inadequate. This is particularly relevant in bridging…

Methodology · Statistics 2025-10-21 Jun Wu , Stéphane Guerrier , Si Gou , Yogeshvar N. Kalia , Luca Insolia

In this paper, we introduce a variation of the group testing problem capturing the idea that a positive test requires a combination of multiple ``types'' of item. Specifically, we assume that there are multiple disjoint \emph{semi-defective…

Information Theory · Computer Science 2024-05-10 Thach V. Bui , Jonathan Scarlett

It is now widely accepted that the standard inferential toolkit used by the scientific research community -- null-hypothesis significance testing (NHST) -- is not fit for purpose. Yet despite the threat posed to the scientific enterprise,…

Methodology · Statistics 2023-11-10 Leonhard Held , Robert Matthews , Manuela Ott , Samuel Pawel

Hypothesis testing is a central statistical method in psychological research and the cognitive sciences. While the problems of null hypothesis significance testing (NHST) have been debated widely, few attractive alternatives exist. In this…

Methodology · Statistics 2020-06-08 Riko Kelter , Julio Michael Stern

Conditional independence (CI) testing is frequently used in data analysis and machine learning for various scientific fields and it forms the basis of constraint-based causal discovery. Oftentimes, CI testing relies on strong, rather…

Methodology · Statistics 2023-06-21 Wiebke Günther , Urmi Ninad , jonas Wahl , Jakob Runge

The aim of this article is to make a contribution to the Bayesian procedure of testing precise hypotheses for parametric models. For this purpose, we define the Bayesian Discrepancy Measure that allows one to evaluate the suitability of a…

Methodology · Statistics 2022-11-21 Francesco Bertolino , Mara Manca , Monica Musio , Walter Racugno , Laura Ventura