English
Related papers

Related papers: Online Experimentation with Surrogate Metrics: Gui…

200 papers

Online controlled experiments, colloquially known as A/B-tests, are the bread and butter of real-world recommender system evaluation. Typically, end-users are randomly assigned some system variant, and a plethora of metrics are then…

Information Retrieval · Computer Science 2024-07-31 Olivier Jeunen , Shubham Baweja , Neeti Pokharna , Aleksei Ustimenko

Online controlled experiments, or A/B tests, are large-scale randomized trials in digital environments. This paper investigates the estimands of the difference-in-means estimator in these experiments, focusing on scenarios with repeated…

Methodology · Statistics 2024-11-12 Sebastian Ankargren , Mattias Frånberg , Mårten Schultzberg

A/B testing has become the cornerstone of decision-making in online markets, guiding how platforms launch new features, optimize pricing strategies, and improve user experience. In practice, we typically employ the pairwise $t$-test to…

Machine Learning · Statistics 2025-10-29 Junpeng Gong , Chunkai Wang , Hao Li , Jinyong Ma , Haoxuan Li , Xu He

Surrogate index approaches have recently become a popular method of estimating longer-term impact from shorter-term outcomes. In this paper, we leverage 1098 test arms from 200 A/B tests at Netflix to empirically investigate to what degree…

Applications · Statistics 2024-02-01 Vickie Zhang , Michael Zhao , and Maria Dimakopoulou , Anh Le , Nathan Kallus

Online controlled experiments are a crucial tool to allow for confident decision-making in technology companies. A North Star metric is defined (such as long-term revenue or user retention), and system variants that statistically…

Machine Learning · Computer Science 2024-06-14 Olivier Jeunen , Aleksei Ustimenko

While there exists a large amount of literature on the general challenges of and best practices for trustworthy online A/B testing, there are limited studies on sample size estimation, which plays a crucial role in trustworthy and efficient…

Methodology · Statistics 2023-08-21 Jing Zhou , Jiannan Lu , Anas Shallah

Surrogate markers are often used in clinical trials to evaluate treatment effects when primary outcomes are costly, invasive, or take a long time to observe. However, reliance on surrogates can lead to the surrogate paradox, where a…

Methodology · Statistics 2025-06-17 Emily Hsiao , Lu Tian , Layla Parast

Online media platforms often need to measure how frequently users are exposed to specific content attributes in order to evaluate trade-offs in A/B experiments. A direct approach is to sample content, label it using a high-quality rubric…

Applications · Statistics 2026-02-19 Zehao Xu , Tony Paek , Kevin O'Sullivan , Attila Dobi

Online experimentation (or A/B testing) has been widely adopted in industry as the gold standard for measuring product impacts. Despite the wide adoption, few literatures discuss A/B testing with quantile metrics. Quantile metrics, such as…

Applications · Statistics 2019-03-22 Min Liu , Xiaohui Sun , Maneesh Varshney , Ya Xu

Online controlled experiments, such as A/B-tests, are commonly used by modern tech companies to enable continuous system improvements. Despite their paramount importance, A/B-tests are expensive: by their very definition, a percentage of…

Machine Learning · Computer Science 2024-01-09 Shubham Baweja , Neeti Pokharna , Aleksei Ustimenko , Olivier Jeunen

Online experiments in internet systems, also known as A/B tests, are used for a wide range of system tuning problems, such as optimizing recommender system ranking policies and learning adaptive streaming controllers. Decision-makers…

Machine Learning · Computer Science 2025-07-01 Qing Feng , Samuel Daulton , Benjamin Letham , Maximilian Balandat , Eytan Bakshy

The Consistency property between surrogate losses and evaluation metrics has been extensively studied to ensure that minimizing a loss leads to metric optimality. However, the direct relationship between different evaluation metrics remains…

Machine Learning · Computer Science 2026-03-10 Yuanhao Pu , Defu Lian , Enhong Chen

Estimating the effects of long-term treatments through A/B testing is challenging. Treatments, such as updates to product functionalities, user interface designs, and recommendation algorithms, are intended to persist within the system for…

Econometrics · Economics 2025-12-30 Shan Huang , Chen Wang , Yuan Yuan , Jinglong Zhao , Brocco , Zhang

A/B tests are randomized experiments frequently used by companies that offer services on the Web for assessing the impact of new features. During an experiment, each user is randomly redirected to one of two versions of the website, called…

Social and Information Networks · Computer Science 2021-08-12 Francisco Galuppo Azevedo , Bruno Demattos Nogueira , Fabricio Murai , Ana Paula Couto da Silva

In the past decade, AB tests have become the standard method for making product decisions in tech companies. They offer a scientific approach to product development, using statistical hypothesis testing to control the risks of incorrect…

Methodology · Statistics 2024-02-20 Mårten Schultzberg , Sebastian Ankargren , Mattias Frånberg

Technology firms conduct randomized controlled experiments ("A/B tests") to learn which actions to take to improve business outcomes. In firms with mature experimentation platforms, experimentation programs can consist of many thousands of…

Methodology · Statistics 2025-05-30 Winston Chou , Colin Gray , Nathan Kallus , Aurélien Bibaut , Simon Ejdemyr

On-line experimentation (also known as A/B testing) has become an integral part of software development. To timely incorporate user feedback and continuously improve products, many software companies have adopted the culture of agile…

Applications · Statistics 2019-08-13 Yu Wang , Somit Gupta , Jiannan Lu , Ali Mahmoudzadeh , Sophia Liu

Online experiments such as Randomised Controlled Trials (RCTs) or A/B-tests are the bread and butter of modern platforms on the web. They are conducted continuously to allow platforms to estimate the causal effect of replacing system…

Machine Learning · Computer Science 2023-04-24 Olivier Jeunen

Computation of document image quality metrics often depends upon the availability of a ground truth image corresponding to the document. This limits the applicability of quality metrics in applications such as hyperparameter optimization of…

Computer Vision and Pattern Recognition · Computer Science 2018-07-10 Prashant Singh , Ekta Vats , Anders Hast

We introduce a dataset comprising commercial machine translations, gathered weekly over six years across 12 translation directions. Since human A/B testing is commonly used, we assume commercial systems improve over time, which enables us…

Computation and Language · Computer Science 2024-10-04 Guojun Wu , Shay B. Cohen , Rico Sennrich
‹ Prev 1 2 3 10 Next ›