English
Related papers

Related papers: A Sharp Test for the Judge Leniency Design

200 papers

Causal inference from observational data is crucial for many disciplines such as medicine and economics. However, sharp bounds for causal effects under relaxations of the unconfoundedness assumption (causal sensitivity analysis) are subject…

Machine Learning · Computer Science 2023-10-17 Dennis Frauen , Valentyn Melnychuk , Stefan Feuerriegel

We consider the extent to which we can learn from a completely randomized experiment whether all individuals have treatment effects that are weakly of the same sign, a condition we call monotonicity. From a classical sampling perspective,…

Econometrics · Economics 2026-01-13 Jiafeng Chen , Jonathan Roth , Jann Spiess

We develop a framework for difference-in-differences designs with staggered treatment adoption and heterogeneous causal effects. We show that conventional regression-based estimators fail to provide unbiased estimates of relevant estimands…

Econometrics · Economics 2024-01-18 Kirill Borusyak , Xavier Jaravel , Jann Spiess

Sensitivity to unmeasured confounding is not typically a primary consideration in designing treated-control comparisons in observational studies. We introduce a framework allowing researchers to optimize robustness to omitted variable bias…

Methodology · Statistics 2024-07-19 Melody Huang , Dan Soriano , Samuel D. Pimentel

Sensitivity analysis informs causal inference by assessing the sensitivity of conclusions to departures from assumptions. The consistency assumption states that there are no hidden versions of treatment and that the outcome arising…

Methodology · Statistics 2025-12-29 Brian Knaeble , Qinyun Lin , Erich Kummerfeld , Kenneth A. Frank

We consider the problem of inference in shift-share research designs. The choice between existing approaches that allow for unrestricted spatial correlation involves tradeoffs, varying in terms of their validity when there are relatively…

Econometrics · Economics 2022-06-03 Luis Alvarez , Bruno Ferman , Raoni Oliveira

Calibration weighting has been widely used to correct selection biases in non-probability sampling, missing data, and causal inference. The main idea is to calibrate the biased sample to the benchmark by adjusting the subject weights.…

Methodology · Statistics 2023-05-30 Chenyin Gao , Shu Yang , Jae Kwang Kim

Reliable certification of Large Language Models (LLMs)-verifying that failure rates are below a safety threshold-is critical yet challenging. While "LLM-as-a-Judge" offers scalability, judge imperfections, noise, and bias can invalidate…

Machine Learning · Computer Science 2026-01-30 Chen Feng , Minghe Shen , Ananth Balashankar , Carsten Gerner-Beuerle , Miguel R. D. Rodrigues

In observational studies of discrimination, the most common statistical approaches consider either the rate at which decisions are made (benchmark tests) or the success rate of those decisions (outcome tests). Both tests, however, have…

Applications · Statistics 2025-03-07 Johann D. Gaebler , Sharad Goel

The safety and robustness of learning-based decision-making systems are under threats from adversarial examples, as imperceptible perturbations can mislead neural networks to completely different outputs. In this paper, we present an…

Machine Learning · Computer Science 2019-11-28 Chao Tang , Yifei Fan , Anthony Yezzi

For large classes of group testing problems, we derive lower bounds for the probability that all significant items are uniquely identified using specially constructed random designs. These bounds allow us to optimize parameters of the…

Statistics Theory · Mathematics 2022-02-17 Jack Noonan , Anatoly Zhigljavsky

We present a randomization-based inferential framework for experiments characterized by a strongly ignorable assignment mechanism where units have independent probabilities of receiving treatment. Previous works on randomization tests often…

Methodology · Statistics 2019-02-01 Zach Branson , Marie-Abele Bind

Offering a promising solution to the scalability challenges associated with human evaluation, the LLM-as-a-judge paradigm is rapidly gaining traction as an approach to evaluating large language models (LLMs). However, there are still many…

Computation and Language · Computer Science 2025-08-19 Aman Singh Thakur , Kartik Choudhary , Venkat Srinik Ramayapally , Sankaran Vaidyanathan , Dieuwke Hupkes

LLM-judged benchmarks are increasingly used to evaluate complex model behaviors, yet their design introduces failure modes absent in conventional ground-truth based benchmarks. We argue that without tight objectives and verifiable…

Machine Learning · Computer Science 2025-10-09 Benjamin Feuer , Chiung-Yi Tseng , Astitwa Sarthak Lathe , Oussama Elachqar , John P Dickerson

The case$^2$ study, also referred to as the case-case study design, is a valuable approach for conducting inference for treatment effects. Unlike traditional case-control studies, the case$^2$ design compares treatment in two types of cases…

Methodology · Statistics 2025-03-03 Kan Chen , Ting Ye , Dylan S. Small

Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, these judges are relied upon to evaluate harmfulness in order to…

Computation and Language · Computer Science 2026-03-17 Leo Schwinn , Moritz Ladenburger , Tim Beyer , Mehrnaz Mofakhami , Gauthier Gidel , Stephan Günnemann

After variable selection, standard inferential procedures for regression parameters may not be uniformly valid; there is no finite-sample size at which a standard test is guaranteed to approximately attain its nominal size. This problem is…

Methodology · Statistics 2020-07-07 Oliver Dukes , Vahe Avagyan , Stijn Vansteelandt

We propose a hypothesis test that allows for many tested restrictions in a heteroskedastic linear regression model. The test compares the conventional F statistic to a critical value that corrects for many restrictions and conditional…

Econometrics · Economics 2023-01-24 Stanislav Anatolyev , Mikkel Sølvsten

Benford's law is often used as a support to critical decisions related to data quality or the presence of data manipulations or even fraud. However, many authors argue that conventional statistical tests will reject the null of data…

Methodology · Statistics 2022-06-16 Roy Cerqueti , Claudio Lupi

The increasing use of deep learning across various domains highlights the importance of understanding the decision-making processes of these black-box models. Recent research focusing on the decision boundaries of deep classifiers, relies…

Machine Learning · Computer Science 2024-08-13 Inês Gomes , Luís F. Teixeira , Jan N. van Rijn , Carlos Soares , André Restivo , Luís Cunha , Moisés Santos