中文
相关论文

相关论文: Test & Roll: Profit-Maximizing A/B Tests

200 篇论文

A/B tests are randomized experiments frequently used by companies that offer services on the Web for assessing the impact of new features. During an experiment, each user is randomly redirected to one of two versions of the website, called…

社会与信息网络 · 计算机科学 2021-08-12 Francisco Galuppo Azevedo , Bruno Demattos Nogueira , Fabricio Murai , Ana Paula Couto da Silva

A/B testing, also known as controlled experiment, bucket testing or splitting testing, has been widely used for evaluating a new feature, service or product in the data-driven decision processes of online websites. The goal of A/B testing…

应用统计 · 统计学 2016-10-26 Bai Jiang , Xiaolin Shi , Hongwei Shang , Zhigeng Geng , Alyssa Glass

AB testing aids business operators with their decision making, and is considered the gold standard method for learning from data to improve digital user experiences. However, there is usually a gap between the requirements of practitioners,…

Online A/B testing plays a critical role in the high-tech industry to guide product development and accelerate innovation. It performs a null hypothesis statistical test to determine which variant is better. However, a typical A/B test…

统计方法学 · 统计学 2021-09-03 Miao Yu , Wenbin Lu , Rui Song

In online randomized experiments or A/B tests, accurate predictions of participant inclusion rates are of paramount importance. These predictions not only guide experimenters in optimizing the experiment's duration but also enhance the…

统计方法学 · 统计学 2024-02-06 Lorenzo Masoero , Mario Beraha , Thomas Richardson , Stefano Favaro

In this paper, we draw attention to a problem that is often overlooked or ignored by companies practicing hypothesis testing (A/B testing) in online environments. We show that conducting experiments on limited inventory that is shared…

概率论 · 数学 2020-06-11 Dennis Bohle , Alexander Marynych , Matthias Meiners

In an A/B test, the typical objective is to measure the total average treatment effect (TATE), which measures the difference between the average outcome if all users were treated and the average outcome if all users were untreated. However,…

应用统计 · 统计学 2020-04-28 David Holtz , Sinan Aral

We propose new, optimal methods for analyzing randomized trials, when it is suspected that treatment effects may differ in two predefined subpopulations. Such sub-populations could be defined by a biomarker or risk factor measured at…

统计方法学 · 统计学 2016-11-26 Michael Rosenblum , Han Liu , and En-Hsu Yen

Experimental testing is vital in the optimization of web applications, and as such A/B testing has been widely adopted as a methodology for determining optimal content for many web applications. While some testing platforms provide…

统计方法学 · 统计学 2017-10-04 Ian E. Fellows

Randomized trials, also known as A/B tests, are used to select between two policies: a control and a treatment. Given a corresponding set of features, we can ideally learn an optimized policy P that maps the A/B test data features to action…

机器学习 · 计算机科学 2018-06-08 Elon Portugaly , Joseph J. Pfeiffer

Online experiments in internet systems, also known as A/B tests, are used for a wide range of system tuning problems, such as optimizing recommender system ranking policies and learning adaptive streaming controllers. Decision-makers…

机器学习 · 计算机科学 2025-07-01 Qing Feng , Samuel Daulton , Benjamin Letham , Maximilian Balandat , Eytan Bakshy

Online controlled experiments, colloquially known as A/B-tests, are the bread and butter of real-world recommender system evaluation. Typically, end-users are randomly assigned some system variant, and a plethora of metrics are then…

信息检索 · 计算机科学 2024-07-31 Olivier Jeunen , Shubham Baweja , Neeti Pokharna , Aleksei Ustimenko

This paper examines how spillover effects in A/B testing can impede organizational progress and develops strategies for mitigating these challenges. We identify a phenomenon termed ``seesaw experimentation'', where a firm's overall…

综合经济学 · 经济学 2025-01-16 Jin Li , Ye Luo , Xiaowei Zhang

Online experimentation, also known as A/B testing, is the gold standard for measuring product impacts and making business decisions in the tech industry. The validity and utility of experiments, however, hinge on unbiasedness and sufficient…

应用统计 · 统计学 2020-12-17 Min Liu , Jialiang Mao , Kang Kang

In clinical studies upon which decisions are based there are two types of errors that can be made: a type I error arises when the decision is taken to declare a positive outcome when the truth is in fact negative, and a type II error arises…

统计方法学 · 统计学 2024-09-19 Andrew P Grieve

Combining test statistics from independent trials or experiments is a popular method of meta-analysis. However, there is very limited theoretical understanding of the power of the combined test, especially in high-dimensional models…

统计理论 · 数学 2023-10-31 Botond Szabó , Aad van der Vaart , Lasse Vuursteen , Harry van Zanten

Adaptive experimental design (AED) methods are increasingly being used in industry as a tool to boost testing throughput or reduce experimentation cost relative to traditional A/B/N testing methods. However, the behavior and guarantees of…

机器学习 · 计算机科学 2024-09-19 Tanner Fiez , Houssam Nassif , Yu-Cheng Chen , Sergio Gamez , Lalit Jain

Technology firms conduct randomized controlled experiments ("A/B tests") to learn which actions to take to improve business outcomes. In firms with mature experimentation platforms, experimentation programs can consist of many thousands of…

统计方法学 · 统计学 2025-05-30 Winston Chou , Colin Gray , Nathan Kallus , Aurélien Bibaut , Simon Ejdemyr

Online controlled experiments, such as A/B-tests, are commonly used by modern tech companies to enable continuous system improvements. Despite their paramount importance, A/B-tests are expensive: by their very definition, a percentage of…

机器学习 · 计算机科学 2024-01-09 Shubham Baweja , Neeti Pokharna , Aleksei Ustimenko , Olivier Jeunen

Large-scale randomized experiments, sometimes called A/B tests, are increasingly prevalent in many industries. Though such experiments are often analyzed via frequentist $t$-tests, arguably such analyses are deficient: $p$-values are hard…

统计方法学 · 统计学 2020-03-27 F. Richard Guo , James McQueen , Thomas S. Richardson