English
Related papers

Related papers: Powerful A/B-Testing Metrics and Where to Find The…

200 papers

Search engines and recommendation systems attempt to continually improve the quality of the experience they afford to their users. Refining the ranker that produces the lists displayed in response to user requests is an important component…

Information Retrieval · Computer Science 2022-06-07 Vishwa Vinay , Manoj Kilaru , David Arbour

The traditional offline approaches are no longer sufficient for building modern recommender systems in domains such as online news services, mainly due to the high dynamics of environment changes and necessity to operate on a large scale…

Information Retrieval · Computer Science 2019-11-26 Joanna Misztal-Radecka , Dominik Rusiecki , Michał Żmuda , Artur Bujak

Digital firms routinely run many online experiments on shared user populations. When product decisions are compositional, such as combinations of interface elements, flows, messages, or incentives, the number of feasible interventions grows…

Machine Learning · Statistics 2026-04-13 Xin Wen , Xi Chen , Will Wei Sun , Yichen Zhang

Two-sided marketplace platforms often run experiments to test the effect of an intervention before launching it platform-wide. A typical approach is to randomize individuals into the treatment group, which receives the intervention, and the…

Methodology · Statistics 2021-04-27 Hannah Li , Geng Zhao , Ramesh Johari , Gabriel Y. Weintraub

Booming in business and a staple analysis in medical trials, the A/B test assesses the effect of an intervention or treatment by comparing its success rate with that of a control condition. Across many practical applications, it is…

Applications · Statistics 2020-11-16 Quentin F. Gronau , K. N. Akash Raj , Eric-Jan Wagenmakers

Platform trials evaluate multiple experimental treatments under a single master protocol, where new treatment arms are added to the trial over time. Given the multiple treatment comparisons, there is the potential for inflation of the…

Methodology · Statistics 2022-02-09 David S. Robertson , James M. S. Wason , Franz König , Martin Posch , Thomas Jaki

Comparing model performances on benchmark datasets is an integral part of measuring and driving progress in artificial intelligence. A model's performance on a benchmark dataset is commonly assessed based on a single or a small set of…

Artificial Intelligence · Computer Science 2021-11-09 Kathrin Blagec , Georg Dorffner , Milad Moradi , Matthias Samwald

Large-scale monitoring, anomaly detection, and root cause analysis of metrics are essential requirements of the internet-services industry. To address the need to continuously monitor millions of metrics, many anomaly detection approaches…

Machine Learning · Computer Science 2022-03-18 Nikhil Galagali

Every design choice will have different effects on different units. However traditional A/B tests are often underpowered to identify these heterogeneous effects. This is especially true when the set of unit-level attributes is…

Artificial Intelligence · Computer Science 2016-11-09 Alexander Peysakhovich , Akos Lada

The selection of the assumed effect size (AES) critically determines the duration of an experiment, and hence its accuracy and efficiency. Traditionally, experimenters determine AES based on domain knowledge. However, this method becomes…

Machine Learning · Computer Science 2025-04-15 Yu Liu , Runzhe Wan , James McQueen , Doug Hains , Jinxiang Gu , Rui Song

This study presents a theoretical analysis on the efficiency of interleaving, an efficient online evaluation method for rankings. Although interleaving has already been applied to production systems, the source of its high efficiency has…

Information Retrieval · Computer Science 2023-06-21 Kojiro Iizuka , Hajime Morita , Makoto P. Kato

The network effect, wherein one user's activity impacts another user, is common in social network platforms. Many new features in social networks are specifically designed to create a network effect, enhancing user engagement. For instance,…

Social and Information Networks · Computer Science 2024-05-22 Wentao Su , Weitao Duan

This paper examines the use of Monte Carlo simulations to understand statistical concepts in A/B testing and Randomized Controlled Trials (RCTs). We discuss the applicability of simulations in understanding false positive rates and estimate…

Applications · Statistics 2024-11-12 Márton Trencséni

The performance of prediction models is often based on "abstract metrics" that estimate the model's ability to limit residual errors between the observed and predicted values. However, meaningful evaluation and selection of prediction…

Machine Learning · Computer Science 2019-05-13 Saima Aman , Yogesh Simmhan , Viktor K. Prasanna

Randomized experiments, or "A/B" tests, remain the gold standard for evaluating the causal effect of a policy intervention or product change. However, experimental settings, such as social networks, where users are interacting and…

Social and Information Networks · Computer Science 2021-02-17 Yuan Yuan , Kristen M. Altenburger , Farshad Kooti

A/B testing is an effective way to assess the potential impacts of two treatments. For A/B tests conducted by IT companies, the test users of A/B testing are often connected and form a social network. The responses of A/B testing can be…

Methodology · Statistics 2023-09-19 Qiong Zhang

During the last decade, the information technology industry has adopted a data-driven culture, relying on online metrics to measure and monitor business performance. Under the setting of big data, the majority of such metrics approximately…

Applications · Statistics 2018-09-14 Alex Deng , Ulf Knoblich , Jiannan Lu

Approaches to recommendation are typically evaluated in one of two ways: (1) via a (simulated) online experiment, often seen as the gold standard, or (2) via some offline evaluation procedure, where the goal is to approximate the outcome of…

Information Retrieval · Computer Science 2024-06-13 Olivier Jeunen , Ivan Potapov , Aleksei Ustimenko

We propose an alternative framework to existing setups for controlling false alarms when multiple A/B tests are run over time. This setup arises in many practical applications, e.g. when pharmaceutical companies test new treatment options…

Machine Learning · Statistics 2017-11-21 Fanny Yang , Aaditya Ramdas , Kevin Jamieson , Martin J. Wainwright

High-dimensional tests are applied to find relevant sets of variables and relevant models. If variables are selected by analyzing the sums of products matrices and a corresponding mean-value test is performed, there is the danger that the…

Methodology · Statistics 2012-02-10 Juergen Laeuter , Maciej Rosolowski , Ekkehard Glimm