English
Related papers

Related papers: Seesaw Experimentation: A/B Tests with Spillovers

200 papers

Motivated by the widespread adoption of iterative project management techniques, we study the effects of workflow -- iterative or sequential -- on innovative behavior and performance. We conduct a series of laboratory experiments. Our first…

General Economics · Economics 2026-03-03 Evgeny Kagan , Christian Jost , Tobias Lieberum , Sebastian Schiffels

Detecting a minor average treatment effect is a major challenge in large-scale applications, where even minimal improvements can have a significant economic impact. Traditional methods, reliant on normal distribution-based or expanded…

Machine Learning · Statistics 2025-07-01 Yu Zhang , Shanshan Zhao , Bokui Wan , Jinjuan Wang , Xiaodong Yan

A/B test, a simple type of controlled experiment, refers to the statistical procedure of experimenting to compare two treatments applied to test subjects. For example, many IT companies frequently conduct A/B tests on their users who are…

Methodology · Statistics 2026-05-12 Qiong Zhang , Lulu Kang

Online controlled experiments are the primary tool for measuring the causal impact of product changes in digital businesses. It is increasingly common for digital products and services to interact with customers in a personalised way. Using…

Methodology · Statistics 2021-07-02 C. H. Bryan Liu , Benjamin Paul Chamberlain

"Spillover" learning is defined as customers' learning about the quality of a service (or product) from their previous experiences with similar yet not identical services. In this paper, we propose a novel, parsimonious and general Bayesian…

Applications · Statistics 2016-07-21 Andrés Musalem , Yan Shang , Jing-Sheng Song

Online experiments (A/B tests) are widely regarded as the gold standard for evaluating recommender system variants and guiding launch decisions. However, a variety of biases can distort the results of the experiment and mislead…

Information Retrieval · Computer Science 2025-09-03 Chen Zheng , Zhenyu Zhao

Software companies have widely used online A/B testing to evaluate the impact of a new technology by offering it to groups of users and comparing it against the unmodified product. However, running online A/B testing needs not only efforts…

Software Engineering · Computer Science 2024-08-12 Jie JW Wu

In online randomized experiments or A/B tests, accurate predictions of participant inclusion rates are of paramount importance. These predictions not only guide experimenters in optimizing the experiment's duration but also enhance the…

Methodology · Statistics 2024-02-06 Lorenzo Masoero , Mario Beraha , Thomas Richardson , Stefano Favaro

Innovations across science and industry are evaluated using randomized trials (a.k.a. A/B tests). While simple and robust, such static designs are inefficient or infeasible for testing many hypotheses. Adaptive designs can greatly improve…

Machine Learning · Computer Science 2024-08-09 Jimmy Wang , Ethan Che , Daniel R. Jiang , Hongseok Namkoong

Bayes factors, in many cases, have been proven to bridge the classic -value based significance testing and bayesian analysis of posterior odds. This paper discusses this phenomena within the binomial A/B testing setup (applicable for…

Other Statistics · Statistics 2019-03-04 Maciej Skorski

Many measurements at collider experiments study physics candidates that are a subset of a collision event. The presence of multiple such candidates in a given event can cause raw biases which are large compared to typical statistical…

High Energy Physics - Experiment · Physics 2019-08-22 Patrick Koppenburg

Language models famously improve under a smooth scaling law, but some specific capabilities exhibit sudden breakthroughs in performance. Advocates of "emergence" view these capabilities as unlocked at a specific scale, but others attribute…

Machine Learning · Computer Science 2026-02-19 Rosie Zhao , Tian Qin , David Alvarez-Melis , Sham Kakade , Naomi Saphra

This article shows how coworker performance affects individual performance evaluation in a teamwork setting at the workplace. We use high-quality data on football matches to measure an important component of individual performance, shooting…

General Economics · Economics 2024-03-25 Enzo Brox , Michael Lechner

This paper studies how to design two-wave experiments in the presence of spillovers for precise inference on treatment effects. We consider units connected through a single network, local dependence among individuals, and a general class of…

Econometrics · Economics 2025-11-25 Davide Viviano

A/B tests have been widely adopted across industries as the golden rule that guides decision making. However, the long-term true north metrics we ultimately want to drive through A/B test may take a long time to mature. In these situations,…

Applications · Statistics 2021-06-04 Weitao Duan , Shan Ba , Chunzhe Zhang

We consider the problem of constructing multiple independent conditional randomization tests using a single dataset. Because the tests are independent, the randomization p-values can be interpreted individually and combined using standard…

Statistics Theory · Mathematics 2024-10-14 Yao Zhang , Qingyuan Zhao

I study identification, estimation and inference for spillover effects in experiments where units' outcomes may depend on the treatment assignments of other units within a group. I show that the commonly-used reduced-form linear-in-means…

Econometrics · Economics 2022-01-21 Gonzalo Vazquez-Bare

Experimentation platforms in industry must often deal with customer trust issues. Platforms must prove the validity of their claims as well as catch issues that arise. As a central quantity estimated by experimentation platforms, the…

Methodology · Statistics 2025-11-21 Kedar Karhadkar , Jack Klys , Daniel Ting , Artem Vorozhtsov , Houssam Nassif

Companies offering web services routinely run randomized online experiments to estimate the causal impact associated with the adoption of new features and policies on key performance metrics of interest. These experiments are used to…

Methodology · Statistics 2023-07-13 Lorenzo Masoero , Doug Hains , James McQueen

Large-scale randomized experiments, sometimes called A/B tests, are increasingly prevalent in many industries. Though such experiments are often analyzed via frequentist $t$-tests, arguably such analyses are deficient: $p$-values are hard…

Methodology · Statistics 2020-03-27 F. Richard Guo , James McQueen , Thomas S. Richardson
‹ Prev 1 3 4 5 6 7 10 Next ›