English
Related papers

Related papers: Evaluating Variance Estimates with Relative Effici…

200 papers

We revisit empirical Bayes discrimination detection, focusing on uncertainty arising from both partial identification and sampling variability. While prior work has mostly focused on partial identification, we find that some empirical…

Econometrics · Economics 2025-08-19 Jiaying Gu , Nikolaos Ignatiadis , Azeem M. Shaikh

Online controlled experiments are a crucial tool to allow for confident decision-making in technology companies. A North Star metric is defined (such as long-term revenue or user retention), and system variants that statistically…

Machine Learning · Computer Science 2024-06-14 Olivier Jeunen , Aleksei Ustimenko

Instrumental variables have been widely used to estimate the causal effect of a treatment on an outcome. Existing confidence intervals for causal effects based on instrumental variables assume that all of the putative instrumental variables…

Methodology · Statistics 2020-06-03 Hyunseung Kang , Youjin Lee , T. Tony Cai , Dylan S. Small

Measuring performance & quantifying a performance change are core evaluation techniques in programming language and systems research. Of 122 recent scientific papers, as many as 65 included experimental evaluation that quantified a…

Methodology · Statistics 2020-07-22 Tomas Kalibera , Richard Jones

Experimentation in online digital platforms is used to inform decision making. Specifically, the goal of many experiments is to optimize a metric of interest. Null hypothesis statistical testing can be ill-suited to this task, as it is…

Methodology · Statistics 2024-12-10 Timothy Sudijono , Simon Ejdemyr , Apoorva Lal , Martin Tingley

We consider the problem of estimating confidence intervals for the mean of a random variable, where the goal is to produce the smallest possible interval for a given number of samples. While minimax optimal algorithms are known for this…

Machine Learning · Statistics 2020-06-19 Shengjia Zhao , Christopher Yeh , Stefano Ermon

This paper tackles the challenge of detecting unreliable behavior in regression algorithms, which may arise from intrinsic variability (e.g., aleatoric uncertainty) or modeling errors (e.g., model uncertainty). First, we formally introduce…

Machine Learning · Computer Science 2024-06-12 Andres Altieri , Marco Romanelli , Georg Pichler , Florence Alberge , Pablo Piantanida

In this paper, we examine the biases that arise when firms run A/B tests on continuous parameters to estimate global treatment effects on performance metrics of interest; we particularly focus on price experiments to measure the price…

Methodology · Statistics 2026-01-22 Ramesh Johari , Orrie B. Page , Gabriel Y. Weintraub

Robustness checks are routine in empirical work, but there is no standard statistical procedure to formally measure what one can learn from them. I propose a "robustness radius" measure to quantify the amount by which the robustness checks…

Econometrics · Economics 2026-02-24 Brenda Prallon

In industry, online randomized controlled experiment (a.k.a. A/B experiment) is a standard approach to measure the impact of a causal change. These experiments have small treatment effect to reduce the potential blast radius. As a result,…

Econometrics · Economics 2025-05-29 Tanmoy Das , Dohyeon Lee , Arnab Sinha

Linear models are foundational tools in statistics and ubiquitous across the applied sciences. However, conventional statistical inference -- such as $t$-tests and $F$-tests -- are only valid at fixed sample sizes, making them unsuitable…

Methodology · Statistics 2025-07-08 Michael Lindon , Dae Woong Ham , Martin Tingley , Iavor Bojinov

In the era of Model-as-a-Service, organizations increasingly rely on third-party AI models for rapid deployment. However, the dynamic nature of emerging AI applications, the continual introduction of new datasets, and the growing number of…

Machine Learning · Computer Science 2026-02-10 Zihan Zhu , Yanqiu Wu , Qiongkai Xu

Multiple hypothesis testing is a fundamental problem in high dimensional inference, with wide applications in many scientific fields. In genome-wide association studies, tens of thousands of tests are performed simultaneously to find if any…

Methodology · Statistics 2011-11-16 Jianqing Fan , Xu Han , Weijie Gu

Some large scale inference problems are considered based on using the relative belief ratio as a measure of statistical evidence. This approach is applied to the multiple testing problem. A particular application of this is concerned with…

Statistics Theory · Mathematics 2016-09-22 Michael Evans , Jabed Tomal

This paper is devoted to testing time series that exhibit behavior related to two or more regimes with different statistical properties. Motivation of our study are two real data sets from plasma physics with observable two-regimes…

Mathematical Physics · Physics 2015-06-04 Janusz gajda , Grzegorz Sikora , Agnieszka Wyłomańska

Stochastic models are widely used to verify whether systems satisfy their reliability, performance and other nonfunctional requirements. However, the validity of the verification depends on how accurately the parameters of these models can…

Software Engineering · Computer Science 2022-02-22 Naif Alasmari , Radu Calinescu , Colin Paterson , Raffaela Mirandola

Parametric statistical methods play a central role in analyzing risk through its underlying frequency and severity components. Given the wide availability of numerical algorithms and high-speed computers, researchers and practitioners often…

Applications · Statistics 2025-06-17 Michael R. Powers , Jiaxin Xu

This paper investigates decision-making in A/B experiments for online platforms and marketplaces. In such settings, due to constraints on inventory, A/B experiments typically lead to biased estimators because of *interference* between…

Methodology · Statistics 2025-08-26 Ramesh Johari , Hannah Li , Anushka Murthy , Gabriel Y. Weintraub

It is difficult to choose detection thresholds for tests of non-stationarity that assume {\em a priori} a noise model if the data is statistically uncharacterized to begin with. This is a potentially serious problem when an automated…

General Relativity and Quantum Cosmology · Physics 2009-12-30 Soumya D. Mohanty

This paper examines the use of Monte Carlo simulations to understand statistical concepts in A/B testing and Randomized Controlled Trials (RCTs). We discuss the applicability of simulations in understanding false positive rates and estimate…

Applications · Statistics 2024-11-12 Márton Trencséni