English
Related papers

Related papers: Active Hypothesis Testing under Computational Budg…

200 papers

Multiple testing of a single hypothesis and testing multiple hypotheses are usually done in terms of p-values. In this paper we replace p-values with their natural competitor, e-values, which are closely related to betting, Bayes factors,…

Statistics Theory · Mathematics 2021-10-26 Vladimir Vovk , Ruodu Wang

We consider the problem of diagnosing faults in a system represented by a Bayesian network, where diagnosis corresponds to recovering the most likely state of unobserved nodes given the outcomes of tests (observed nodes). Finding an optimal…

Artificial Intelligence · Computer Science 2012-07-09 Alice X. Zheng , Irina Rish , Alina Beygelzimer

Human annotation cost and time remain significant bottlenecks in Natural Language Processing (NLP), with test data annotation being particularly expensive due to the stringent requirement for low-error and high-quality labels necessary for…

Computation and Language · Computer Science 2026-03-24 Antonio Purificato , Maria Sofia Bucarelli , Andrea Bacciu , Amin Mantrach , Fabrizio Silvestri

Hypothesis testing is a statistical method used to draw conclusions about populations from sample data, typically represented in tables. With the prevalence of graph representations in real-life applications, hypothesis testing in graphs is…

Machine Learning · Statistics 2025-02-27 Yun Wang , Chrysanthi Kosyfaki , Sihem Amer-Yahia , Reynold Cheng

Consider the problem of testing whether the outputs of a large language model (LLM) system change under an arbitrary intervention, such as an input perturbation or changing the model variant. We cannot simply compare two LLM outputs since…

Computation and Language · Computer Science 2025-06-10 Paulius Rauba , Qiyao Wei , Mihaela van der Schaar

In many settings, robust data analysis involves computational methods for uncertainty quantification and statistical inference. To design frequentist studies that leverage robust analysis methods, suitable sample sizes to achieve desired…

Methodology · Statistics 2025-12-19 Luke Hagar , Andrew J. Martin

Much of science is (rightly or wrongly) driven by hypothesis testing. Even in situations where the hypothesis testing paradigm is correct, the common practice of basing inferences solely on p-values has been under intense criticism for over…

Methodology · Statistics 2015-12-31 M. J. Bayarri , Daniel J. Benjamin , James O. Berger , Thomas M. Sellke

Recently, a new testing approach for response-adaptive clinical trials was proposed based on the allocation probabilities (AP) rather than the outcome data. While original work on the AP test focused on binary and normal endpoints and…

Methodology · Statistics 2026-05-11 Stina Zetterstrom , David S. Robertson , Thomas Jaki , Sofía S. Villar

"Asymptotic formulae for likelihood-based tests of new physics" presents a mathematical formalism for a new approximation for hypothesis testing in high energy physics. The approximations are designed to greatly reduce the computational…

High Energy Physics - Experiment · Physics 2011-10-25 Eric Burns , Wade Fisher

Compared to p-values, e-values provably guarantee safe, valid inference. If the goal is to test multiple hypotheses simultaneously, one can construct e-values for each individual test and then use the recently developed e-BH procedure to…

Methodology · Statistics 2024-12-03 Neil Dey , Ryan Martin , Jonathan P. Williams

Adaptive experiments use preliminary analyses of the data to inform further course of action and are commonly used in many disciplines including medical and social sciences. Because the null hypothesis and experimental design are…

Methodology · Statistics 2026-05-26 Tobias Freidling , Qingyuan Zhao , Zijun Gao

Test-time compute scaling, the practice of spending extra computation during inference via repeated sampling, search, or extended reasoning, has become a powerful lever for improving large language model performance. Yet deploying these…

Machine Learning · Computer Science 2026-04-17 Zhiyuan Zhai , Bingcong Li , Bingnan Xiao , Ming Li , Xin Wang

The problem of handling adaptivity in data analysis, intentional or not, permeates a variety of fields, including test-set overfitting in ML challenges and the accumulation of invalid scientific discoveries. We propose a mechanism for…

Machine Learning · Computer Science 2019-04-03 Blake Woodworth , Vitaly Feldman , Saharon Rosset , Nathan Srebro

The self-rationalising capabilities of LLMs are appealing because the generated explanations can give insights into the plausibility of the predictions. However, how faithful the explanations are to the predictions is questionable, raising…

Computation and Language · Computer Science 2024-12-18 Marc Braun , Jenny Kunz

We are concerned with testing replicability hypotheses for many endpoints simultaneously. This constitutes a multiple test problem with composite null hypotheses. Traditional $p$-values, which are computed under least favourable parameter…

Methodology · Statistics 2020-02-26 Anh-Tuan Hoang , Thorsten Dickhaus

Test-time computation has become a primary driver of progress in large language model (LLM) reasoning, but it is increasingly bottlenecked by expensive verification. In many reasoning systems, a large fraction of verifier calls are spent on…

Artificial Intelligence · Computer Science 2026-02-05 Shuhui Qu

Active testing enables label-efficient evaluation of predictive models through careful data acquisition, but it can pose a significant computational cost. We identify cost-saving measures that enable active testing to be scaled up to large…

Machine Learning · Computer Science 2025-11-26 Gabrielle Berrada , Jannik Kossen , Freddie Bickford Smith , Muhammed Razzak , Yarin Gal , Tom Rainforth

Empirical research in Natural Language Processing (NLP) has adopted a narrow set of principles for assessing hypotheses, relying mainly on p-value computation, which suffers from several known issues. While alternative proposals have been…

Computation and Language · Computer Science 2020-05-06 Erfan Sadeqi Azer , Daniel Khashabi , Ashish Sabharwal , Dan Roth

The randomized $p$-value, (nonrandomized) mid-$p$-value and abstract randomized $p$-value have all been recommended for testing a null hypothesis whenever the test statistic has a discrete distribution. This paper provides a unifying…

Computation · Statistics 2014-12-02 Joshua D Habiger

We examine hypothesis testing within a principal-agent framework, where a strategic agent, holding private beliefs about the effectiveness of a product, submits data to a principal who decides on approval. The principal employs a hypothesis…

Machine Learning · Computer Science 2025-08-06 Safwan Hossain , Yatong Chen , Yiling Chen