中文
相关论文

相关论文: Powerful A/B-Testing Metrics and Where to Find The…

200 篇论文

Search engines and recommendation systems attempt to continually improve the quality of the experience they afford to their users. Refining the ranker that produces the lists displayed in response to user requests is an important component…

信息检索 · 计算机科学 2022-06-07 Vishwa Vinay , Manoj Kilaru , David Arbour

The traditional offline approaches are no longer sufficient for building modern recommender systems in domains such as online news services, mainly due to the high dynamics of environment changes and necessity to operate on a large scale…

信息检索 · 计算机科学 2019-11-26 Joanna Misztal-Radecka , Dominik Rusiecki , Michał Żmuda , Artur Bujak

Digital firms routinely run many online experiments on shared user populations. When product decisions are compositional, such as combinations of interface elements, flows, messages, or incentives, the number of feasible interventions grows…

机器学习 · 统计学 2026-04-13 Xin Wen , Xi Chen , Will Wei Sun , Yichen Zhang

Two-sided marketplace platforms often run experiments to test the effect of an intervention before launching it platform-wide. A typical approach is to randomize individuals into the treatment group, which receives the intervention, and the…

统计方法学 · 统计学 2021-04-27 Hannah Li , Geng Zhao , Ramesh Johari , Gabriel Y. Weintraub

Booming in business and a staple analysis in medical trials, the A/B test assesses the effect of an intervention or treatment by comparing its success rate with that of a control condition. Across many practical applications, it is…

应用统计 · 统计学 2020-11-16 Quentin F. Gronau , K. N. Akash Raj , Eric-Jan Wagenmakers

Platform trials evaluate multiple experimental treatments under a single master protocol, where new treatment arms are added to the trial over time. Given the multiple treatment comparisons, there is the potential for inflation of the…

统计方法学 · 统计学 2022-02-09 David S. Robertson , James M. S. Wason , Franz König , Martin Posch , Thomas Jaki

Comparing model performances on benchmark datasets is an integral part of measuring and driving progress in artificial intelligence. A model's performance on a benchmark dataset is commonly assessed based on a single or a small set of…

人工智能 · 计算机科学 2021-11-09 Kathrin Blagec , Georg Dorffner , Milad Moradi , Matthias Samwald

Large-scale monitoring, anomaly detection, and root cause analysis of metrics are essential requirements of the internet-services industry. To address the need to continuously monitor millions of metrics, many anomaly detection approaches…

机器学习 · 计算机科学 2022-03-18 Nikhil Galagali

Every design choice will have different effects on different units. However traditional A/B tests are often underpowered to identify these heterogeneous effects. This is especially true when the set of unit-level attributes is…

人工智能 · 计算机科学 2016-11-09 Alexander Peysakhovich , Akos Lada

The selection of the assumed effect size (AES) critically determines the duration of an experiment, and hence its accuracy and efficiency. Traditionally, experimenters determine AES based on domain knowledge. However, this method becomes…

机器学习 · 计算机科学 2025-04-15 Yu Liu , Runzhe Wan , James McQueen , Doug Hains , Jinxiang Gu , Rui Song

This study presents a theoretical analysis on the efficiency of interleaving, an efficient online evaluation method for rankings. Although interleaving has already been applied to production systems, the source of its high efficiency has…

信息检索 · 计算机科学 2023-06-21 Kojiro Iizuka , Hajime Morita , Makoto P. Kato

The network effect, wherein one user's activity impacts another user, is common in social network platforms. Many new features in social networks are specifically designed to create a network effect, enhancing user engagement. For instance,…

社会与信息网络 · 计算机科学 2024-05-22 Wentao Su , Weitao Duan

This paper examines the use of Monte Carlo simulations to understand statistical concepts in A/B testing and Randomized Controlled Trials (RCTs). We discuss the applicability of simulations in understanding false positive rates and estimate…

应用统计 · 统计学 2024-11-12 Márton Trencséni

The performance of prediction models is often based on "abstract metrics" that estimate the model's ability to limit residual errors between the observed and predicted values. However, meaningful evaluation and selection of prediction…

机器学习 · 计算机科学 2019-05-13 Saima Aman , Yogesh Simmhan , Viktor K. Prasanna

Randomized experiments, or "A/B" tests, remain the gold standard for evaluating the causal effect of a policy intervention or product change. However, experimental settings, such as social networks, where users are interacting and…

社会与信息网络 · 计算机科学 2021-02-17 Yuan Yuan , Kristen M. Altenburger , Farshad Kooti

A/B testing is an effective way to assess the potential impacts of two treatments. For A/B tests conducted by IT companies, the test users of A/B testing are often connected and form a social network. The responses of A/B testing can be…

统计方法学 · 统计学 2023-09-19 Qiong Zhang

During the last decade, the information technology industry has adopted a data-driven culture, relying on online metrics to measure and monitor business performance. Under the setting of big data, the majority of such metrics approximately…

应用统计 · 统计学 2018-09-14 Alex Deng , Ulf Knoblich , Jiannan Lu

Approaches to recommendation are typically evaluated in one of two ways: (1) via a (simulated) online experiment, often seen as the gold standard, or (2) via some offline evaluation procedure, where the goal is to approximate the outcome of…

信息检索 · 计算机科学 2024-06-13 Olivier Jeunen , Ivan Potapov , Aleksei Ustimenko

We propose an alternative framework to existing setups for controlling false alarms when multiple A/B tests are run over time. This setup arises in many practical applications, e.g. when pharmaceutical companies test new treatment options…

机器学习 · 统计学 2017-11-21 Fanny Yang , Aaditya Ramdas , Kevin Jamieson , Martin J. Wainwright

High-dimensional tests are applied to find relevant sets of variables and relevant models. If variables are selected by analyzing the sums of products matrices and a corresponding mean-value test is performed, there is the danger that the…

统计方法学 · 统计学 2012-02-10 Juergen Laeuter , Maciej Rosolowski , Ekkehard Glimm