English
Related papers

Related papers: Improved Tests for Mediation

200 papers

We revisit the fundamental question of simple-versus-simple hypothesis testing with an eye towards computational complexity, as the statistically optimal likelihood ratio test is often computationally intractable in high-dimensional…

Statistics Theory · Mathematics 2025-05-05 Ankur Moitra , Alexander S. Wein

Many scientific applications involve testing theories that are only partially specified. This task often amounts to testing the goodness-of-fit of a candidate distribution while allowing for reasonable deviations from it. The tolerant…

Statistics Theory · Mathematics 2026-01-28 Lucas Kania , Tudor Manole , Larry Wasserman , Sivaraman Balakrishnan

The asymptotic efficiency, ARE_{p,2}, of the tests for multivariate means theta in \R^d based on the p-means relative to the standard 2-mean, (approximate) likelihood ratio test (LRT), is considered for large dimensions d. It turns out that…

Statistics Theory · Mathematics 2010-06-04 Iosif Pinelis

Mediation hypothesis testing for a large number of mediators is challenging due to the composite structure of the null hypothesis, H0:alpha*beta=0 (alpha: effect of the exposure on the mediator after adjusting for confounders; beta: effect…

Methodology · Statistics 2022-03-28 Jiacong Du , Xiang Zhou , Wei Hao , Yongmei Liu , Jennifer A. Smith , Bhramar Mukherjee

This article introduces a robust hypothesis testing procedure: the Lq-likelihood-ratio-type test (LqRT). By deriving the asymptotic distribution of this test statistic, the authors demonstrate its robustness both analytically and…

Applications · Statistics 2016-09-27 Yichen Qin , Carey E. Priebe

New Large Language Models (LLMs) become available every few weeks, and modern application developers confronted with the unenviable task of having to decide if they should switch to a new model. While human evaluation remains the gold…

Artificial Intelligence · Computer Science 2025-12-25 Suryaansh Jain , Umair Z. Ahmed , Shubham Sahai , Ben Leong

For random samples of size n obtained from p-variate normal distributions, we consider the classical likelihood ratio tests (LRT) for their means and covariance matrices in the high-dimensional setting. These test statistics have been…

Statistics Theory · Mathematics 2013-06-04 Tiefeng Jiang , Fan Yang

Problem Roulette (PR), an online study service at the University of Michigan, offers points-free formative practice to students preparing for examinations in introductory STEM courses. Using four years of PR data involving millions of…

Physics Education · Physics 2023-05-02 William K. Black , Rebecca L. Matz , Mark Mills , A. E. Evrard

We ask: Can focusing on likely classes of a single, in-domain sample improve model predictions? Prior work argued ``no''. We put forward a novel rationale in favor of ``yes'': Sharedness of features among classes indicates their reliability…

Machine Learning · Computer Science 2025-12-23 Johannes Schneider

The Birnbaum-Saunders regression model is commonly used in reliability studies. We address the issue of performing inference in this class of models when the number of observations is small. We show that the likelihood ratio test tends to…

Methodology · Statistics 2009-11-25 Artur J. Lemonte , Silvia L. P. Ferrari , Francisco Cribari-Neto

Threshold tests have recently been proposed as a useful method for detecting bias in lending, hiring, and policing decisions. For example, in the case of credit extensions, these tests aim to estimate the bar for granting loans to white and…

Machine Learning · Statistics 2018-03-13 Emma Pierson , Sam Corbett-Davies , Sharad Goel

Multimodal large language models (MLLMs) hold considerable promise for applications in healthcare. However, their deployment in safety-critical settings is hindered by two key limitations: (i) sensitivity to prompt design, and (ii) a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Anita Kriz , Elizabeth Laura Janes , Xing Shen , Tal Arbel

Mediation analysis in high-dimensional settings often involves identifying potential mediators among a large number of measured variables. For this purpose, a two-step familywise error rate procedure called ScreenMin has been recently…

Methodology · Statistics 2020-07-07 Vera Djordjilović , Jesse Hemerik , Magne Thoresen

We propose a new class of weighted logrank tests (WLRT) that control the risk of concluding that a new drug is more efficacious than standard of care, when, in fact, it is uniformly inferior. Perhaps surprisingly, this risk is not…

Applications · Statistics 2018-07-31 Dominic Magirr , Carl-Fredrik Burman

Although large language models (LLMs) are increasingly trained using human feedback for safety and alignment with human values, alignment decisions often overlook human social diversity. This study examines how incorporating pluralistic…

Artificial Intelligence · Computer Science 2025-11-27 Dalia Ali , Dora Zhao , Allison Koenecke , Orestis Papakyriakopoulos

We introduce a new procedure for testing the significance of a set of regression coefficients in a Gaussian linear model with $n \geq d$. Our method, the $L$-test, provides the same statistical validity guarantee as the classical $F$-test,…

Methodology · Statistics 2025-12-01 Danielle Paulson , Souhardya Sengupta , Lucas Janson

This paper investigates improved testing inferences under a general multivariate elliptical regression model. The model is very flexible in terms of the specification of the mean vector and the dispersion matrix, and of the choice of the…

Statistics Theory · Mathematics 2016-11-01 T. F. N. Melo , S. L. P. Ferrari , A. G. Patriota

Multi-agent LLM debate improves factuality and reasoning, but most recipes pick a fixed round count, over-spending on easy items and under-spending on hard ones. We adapt Wald's Sequential Probability Ratio Test (SPRT) as a plug-in compute…

Machine Learning · Computer Science 2026-05-20 Andrea Morandi

Mature test automation is key for achieving software quality at speed. In this paper, we present a multivocal literature review with the objective to survey and synthesize the guidelines given in the literature for improving test automation…

Software Engineering · Computer Science 2022-02-21 Yuqing Wang , Mika V. Mäntylä , Zihao Liu , Jouni Markkula , Päivi Raulamo-jurvanen

Inference in models where the parameter is defined by moment inequalities is of interest in many areas of economics. This paper develops a new method for improving the performance of generalized moment selection (GMS) testing procedures in…

Econometrics · Economics 2020-08-26 Rami V. Tabri , Christopher D. Walker