English
Related papers

Related papers: Enhancing External Validity of Experiments with On…

200 papers

Continuous integration at scale is costly but essential to software development. Various test optimization techniques including test selection and prioritization aim to reduce the cost. Test batching is an effective alternative, but…

Software Engineering · Computer Science 2023-08-28 Emad Fallahzadeh , Amir Hossein Bavand , Peter C. Rigby

Randomized experiments are an excellent tool for estimating internally valid causal effects with the sample at hand, but their external validity is frequently debated. While classical results on the estimation of Population Average…

Methodology · Statistics 2023-01-13 Apoorva Lal , Wenjing Zheng , Simon Ejdemyr

Online experiments such as Randomised Controlled Trials (RCTs) or A/B-tests are the bread and butter of modern platforms on the web. They are conducted continuously to allow platforms to estimate the causal effect of replacing system…

Machine Learning · Computer Science 2023-04-24 Olivier Jeunen

Online controlled experiments, or A/B tests, are large-scale randomized trials in digital environments. This paper investigates the estimands of the difference-in-means estimator in these experiments, focusing on scenarios with repeated…

Methodology · Statistics 2024-11-12 Sebastian Ankargren , Mattias Frånberg , Mårten Schultzberg

Controlled experimentation, also called A/B testing, is widely adopted to accelerate product innovations in the online world. However, how fast we innovate can be limited by how we run experiments. Most experiments go through a "ramp up"…

Applications · Statistics 2018-01-26 Ya Xu , Weitao Duan , Shaochen Huang

Randomized experiments ensure robust causal inference that are critical to effective learning analytics research and practice. However, traditional randomized experiments, like A/B tests, are limiting in large scale digital learning…

Applications · Statistics 2019-02-04 Timothy NeCamp , Josh Gardner , Christopher Brooks

Online experimentation (or A/B testing) has been widely adopted in industry as the gold standard for measuring product impacts. Despite the wide adoption, few literatures discuss A/B testing with quantile metrics. Quantile metrics, such as…

Applications · Statistics 2019-03-22 Min Liu , Xiaohui Sun , Maneesh Varshney , Ya Xu

Reliable estimation of treatment effects from observational data is important in many disciplines such as medicine. However, estimation is challenging when unconfoundedness as a standard assumption in the causal inference literature is…

Machine Learning · Computer Science 2024-10-15 Jonas Schweisthal , Dennis Frauen , Maresa Schröder , Konstantin Hess , Niki Kilbertus , Stefan Feuerriegel

The popularity of online surveys has increased the prominence of using weights that capture units' probabilities of inclusion for claims of representativeness. Yet, much uncertainty remains regarding how these weights should be employed in…

Methodology · Statistics 2017-08-16 Luke W. Miratrix , Jasjeet S. Sekhon , Alexander G. Theodoridis , Luis F. Campos

We present unexpected findings from a large-scale benchmark study evaluating Conditional Average Treatment Effect (CATE) estimation algorithms, i.e., CATE models. By running 16 modern CATE models on 12 datasets and 43,200 sampled variants…

Machine Learning · Statistics 2025-02-21 Haining Yu , Yizhou Sun

Online A/B tests have become increasingly popular and important for social platforms. However, accurately estimating the global average treatment effect (GATE) has proven to be challenging due to network interference, which violates the…

Methodology · Statistics 2023-11-27 Qianyi Chen , Bo Li , Lu Deng , Yong Wang

A/B testing is the foundation of decision-making in online platforms, yet social products often suffer from network interference: user interactions cause treatment effects to spill over into the control group. Such spillovers bias causal…

Social and Information Networks · Computer Science 2026-02-10 Xu Min , Zhaoxu Yang , Kaixuan Tan , Juan Yan , Xunbin Xiong , Zihao Zhu , Kaiyu Zhu , Fenglin Cui , Yang Yang , Sihua Yang , Jianhui Bu

With the advancement in technology, raw event data generated by the digital world have grown tremendously. However, such data tend to be insufficient and noisy when it comes to measuring user intention or satisfaction. One effective way to…

Applications · Statistics 2019-06-25 Weitao Duan , Qian Wang , Rogier Verhulst , Ya Xu

Online controlled experiments (A/B tests) have become the gold standard for learning the impact of new product features in technology companies. Randomization enables the inference of causality from an A/B test. The randomized assignment…

Applications · Statistics 2022-12-20 Qike Li , Samir Jamkhande , Pavel Kochetkov , Pai Liu

A/B testing is critical for modern technological companies to evaluate the effectiveness of newly developed products against standard baselines. This paper studies optimal designs that aim to maximize the amount of information obtained from…

Methodology · Statistics 2023-11-07 Ting Li , Chengchun Shi , Jianing Wang , Fan Zhou , Hongtu Zhu

Recently, many causal estimators for Conditional Average Treatment Effect (CATE) and instrumental variable (IV) problems have been published and open sourced, allowing to estimate granular impact of both randomized treatments (such as A/B…

Machine Learning · Computer Science 2022-12-21 Egor Kraev , Timo Flesch , Hudson Taylor Lekunze , Mark Harley , Pere Planell Morell

Online A/B testing plays a critical role in the high-tech industry to guide product development and accelerate innovation. It performs a null hypothesis statistical test to determine which variant is better. However, a typical A/B test…

Methodology · Statistics 2021-09-03 Miao Yu , Wenbin Lu , Rui Song

Two-phase sampling is a simple and cost-effective estimation strategy in survey sampling and is widely used in practice. Because the phase-2 sampling probability typically depends on low-cost variables collected at phase 1, naive estimation…

Methodology · Statistics 2025-11-11 Kazuharu Harada , Masataka Taguri

Large-scale randomized experiments, sometimes called A/B tests, are increasingly prevalent in many industries. Though such experiments are often analyzed via frequentist $t$-tests, arguably such analyses are deficient: $p$-values are hard…

Methodology · Statistics 2020-03-27 F. Richard Guo , James McQueen , Thomas S. Richardson

In streaming platforms churn is extremely costly, yet A/B tests are typically evaluated using outcomes observed within a limited experimental horizon. Even when both short- and predicted long-term engagement metrics are considered, they may…

Machine Learning · Computer Science 2026-04-23 Dario Simionato , Andrea Tonon , Mingxue Wang , Weiguo Wang , Tong Gui , Xiaoyue Li