English
Related papers

Related papers: REACT to NHST: Sensible conclusions to meaningful …

200 papers

Many studies include a goal of determining whether there is treatment effect heterogeneity across different subpopulations. In this paper, we propose a U-statistic-based non-parametric test of the null hypothesis that the treatment effects…

Methodology · Statistics 2020-12-08 Maozhu Dai , Hal S. Stern

Large language model editing methods frequently suffer from overfitting, wherein factual updates can propagate beyond their intended scope, overemphasizing the edited target even when it's contextually inappropriate. To address this…

Artificial Intelligence · Computer Science 2025-05-27 Haitian Zhong , Yuhuan Liu , Ziyang Xu , Guofan Liu , Qiang Liu , Shu Wu , Zhe Zhao , Liang Wang , Tieniu Tan

The objective of this paper is to develop statistical methodology for planning and evaluating three-armed non-inferiority trials for general retention of effect hypotheses, where the endpoint of interest may follow any (regular) parametric…

Methodology · Statistics 2009-12-22 M. Mielke , A. Munk

This article introduces a robust hypothesis testing procedure: the Lq-likelihood-ratio-type test (LqRT). By deriving the asymptotic distribution of this test statistic, the authors demonstrate its robustness both analytically and…

Applications · Statistics 2016-09-27 Yichen Qin , Carey E. Priebe

We conducted a systematic comparison of statistical methods used for the analysis of time-to-event outcomes under various proportional and nonproportional hazard (NPH) scenarios. Our study used data from recently published oncology trials…

Applications · Statistics 2025-02-12 Xinyu Zhang , Erich J. Greene , Ondrej Blaha , Wei Wei

Statistical significance testing plays an important role when drawing conclusions from experimental results in NLP papers. Particularly, it is a valuable tool when one would like to establish the superiority of one algorithm over another.…

Computation and Language · Computer Science 2018-09-06 Rotem Dror , Roi Reichart

Two-sample testing is a fundamental problem in statistics, and many famous two-sample tests are designed to be fully non-parametric. These existing methods perform well with location and scale shifts but are less robust when faced with more…

Methodology · Statistics 2021-10-12 Benjamin Brown , Kai Zhang

Dense retrieval is a basic building block of information retrieval applications. One of the main challenges of dense retrieval in real-world settings is the handling of queries containing misspelled words. A popular approach for handling…

Experiments often yield non-identically distributed data for statistical analysis. Tests of hypothesis under such set-ups are generally performed using the likelihood ratio test, which is non-robust with respect to outliers and model…

Statistics Theory · Mathematics 2017-07-25 Abhik Ghosh , Ayanendranath Basu

An important aspect of multiple hypothesis testing is controlling the significance level, or the level of Type I error. When the test statistics are not independent it can be particularly challenging to deal with this problem, without…

Statistics Theory · Mathematics 2009-03-04 Sandy Clarke , Peter Hall

Randomized Controlled Trials (RCTs) represent a gold standard when developing policy guidelines. However, RCTs are often narrow, and lack data on broader populations of interest. Causal effects in these populations are often estimated using…

Machine Learning · Computer Science 2023-03-07 Zeshan Hussain , Michael Oberst , Ming-Chieh Shih , David Sontag

Many researchers have applied classical statistical decision theory to evaluate treatment choices and learn optimal policies. However, because this framework is based solely on realized outcomes under chosen decisions and ignores…

Statistics Theory · Mathematics 2025-10-21 Benedikt Koch , Kosuke Imai

[See paper for full abstract] Meta-analysis is a crucial tool for answering scientific questions. It is usually conducted on a relatively small amount of ``trusted'' data -- ideally from randomized, controlled trials -- which allow causal…

Machine Learning · Statistics 2024-07-15 Shiva Kaul , Geoffrey J. Gordon

In modern scientific research, small-scale studies with limited participants are increasingly common. However, interpreting individual outcomes can be challenging, making it standard practice to combine data across studies using random…

Statistics Theory · Mathematics 2025-11-04 Lucas Kania , Larry Wasserman , Sivaraman Balakrishnan

Searches for unknown physics and decisions between competing astrophysical models to explain data both rely on statistical hypothesis testing. The usual approach in searches for new physical phenomena is based on the statistical Likelihood…

Data Analysis, Statistics and Probability · Physics 2016-02-22 Sara Algeri , Jan Conrad , David A. van Dyk

The mechanisms underlying scientific confabulation in Large Language Models (LLMs) remain poorly understood. We introduce ReFACT (Reddit False And Correct Texts), a benchmark of 1,001 expert-annotated question-answer pairs with span-level…

Computation and Language · Computer Science 2026-04-24 Yindong Wang , Martin Preiß , Margarita Bugueño , Jan Vincent Hoffbauer , Abdullatif Ghajar , Tolga Buz , Gerard de Melo

Understanding causality should be a core requirement of any attempt to build real impact through AI. Due to the inherent unobservability of counterfactuals, large randomised trials (RCTs) are the standard for causal inference. But large…

The standard paradigm for confirmatory clinical trials is to compare experimental treatments with a control, for example the standard of care or a placebo. However, it is not always the case that a suitable control exists. Efficient…

Methodology · Statistics 2024-10-29 Thomas Burnett , Thomas Jaki

A large fraction of papers in the climate literature includes erroneous uses of significance tests. A Bayesian analysis is presented to highlight the meaning of significance tests and why typical misuse occurs. It is concluded that a…

Atmospheric and Oceanic Physics · Physics 2016-08-24 Maarten H. P. Ambaum

Recent Large Reasoning Models (LRMs) have achieved remarkable progress on task-specific benchmarks, yet their evaluation methods remain constrained by isolated problem-solving paradigms. Existing benchmarks predominantly assess…

Computation and Language · Computer Science 2025-07-16 Zhuoshi Pan , Qizhi Pei , Yu Li , Qiyao Sun , Zinan Tang , H. Vicky Zhao , Conghui He , Lijun Wu
‹ Prev 1 8 9 10 Next ›