中文
相关论文

相关论文: Powerful A/B-Testing Metrics and Where to Find The…

200 篇论文

In the past decade, AB tests have become the standard method for making product decisions in tech companies. They offer a scientific approach to product development, using statistical hypothesis testing to control the risks of incorrect…

统计方法学 · 统计学 2024-02-20 Mårten Schultzberg , Sebastian Ankargren , Mattias Frånberg

Businesses frequently run online controlled experiments (i.e., A/B tests) to learn about the effect of an intervention on multiple business metrics. To account for multiple hypothesis testing, multiple metrics are commonly aggregated into a…

统计方法学 · 统计学 2026-01-22 Luke Hagar , Nathaniel T. Stevens

In this paper, we address the fundamental statistical question: how can you assess the power of an A/B test when the units in the study are exposed to interference? This question is germane to many scientific and industrial practitioners…

社会与信息网络 · 计算机科学 2017-10-12 James D. Wilson , David T. Uminsky

In this paper, we present our work towards comparing on-line and off-line evaluation metrics in the context of small e-commerce recommender systems. Recommending on small e-commerce enterprises is rather challenging due to the lower volume…

信息检索 · 计算机科学 2020-06-11 Ladislav Peska , Peter Vojtas

A/B testing refers to the statistical procedure of conducting an experiment to compare two treatments, A and B, applied to different testing subjects. It is widely used by technology companies such as Facebook, LinkedIn, and Netflix, to…

统计方法学 · 统计学 2026-05-12 Victoria Pokhiko , Qiong Zhang , Lulu Kang , D'arcy P. Mays

This paper investigates decision-making in A/B experiments for online platforms and marketplaces. In such settings, due to constraints on inventory, A/B experiments typically lead to biased estimators because of *interference* between…

统计方法学 · 统计学 2025-08-26 Ramesh Johari , Hannah Li , Anushka Murthy , Gabriel Y. Weintraub

Tech companies (e.g., Google or Facebook) often use randomized online experiments and/or A/B testing primarily based on the average treatment effects to compare their new product with an old one. However, it is also critically important to…

统计方法学 · 统计学 2021-11-09 Chengchun Shi , Shikai Luo , Hongtu Zhu , Rui Song

In this paper, we provide a statistical testing framework to check whether a random sample splitting in a multi-dimensional space is carried out in a valid way, which could be directly applied to A/B testing and multivariate testing to…

统计方法学 · 统计学 2018-10-11 Jing Miao , Hongyuan Yuan , Zhenyu Yan

We develop a theoretical framework for sample splitting in A/B testing environments, where data for each test are partitioned into two splits to measure methodological performance when the true impacts of tests are unobserved. We show that…

计量经济学 · 经济学 2026-03-24 Ryan Kessler , James McQueen , Miikka Rokkanen

E-commerce companies have a number of online products, such as organic search, sponsored search, and recommendation modules, to fulfill customer needs. Although each of these products provides a unique opportunity for users to interact with…

应用统计 · 统计学 2020-06-23 Xuan Yin , Liangjie Hong

Experimentation platforms in industry must often deal with customer trust issues. Platforms must prove the validity of their claims as well as catch issues that arise. As a central quantity estimated by experimentation platforms, the…

统计方法学 · 统计学 2025-11-21 Kedar Karhadkar , Jack Klys , Daniel Ting , Artem Vorozhtsov , Houssam Nassif

Real-world recommender systems often need to balance multiple objectives when deciding which recommendations to present to users. These include behavioural signals (e.g. clicks, shares, dwell time), as well as broader objectives (e.g.…

信息检索 · 计算机科学 2024-09-17 Olivier Jeunen , Jatin Mandav , Ivan Potapov , Nakul Agarwal , Sourabh Vaid , Wenzhe Shi , Aleksei Ustimenko

Though it has been recognized that recommending serendipitous (i.e., surprising and relevant) items can be helpful for increasing users' satisfaction and behavioral intention, how to measure serendipity in the offline environment is still…

人机交互 · 计算机科学 2020-04-23 Li Chen , Ningxia Wang , Yonghua Yang , Keping Yang , Quan Yuan

eBay's experimentation platform runs hundreds of A/B tests on any given day. The platform integrates with the tracking infrastructure and customer experience servers, provides the sampling service for experiments, and has the responsibility…

应用统计 · 统计学 2023-03-10 Keyu Nie , Zezhong Zhang , Bingquan Xu , Tao Yuan

A/B-tests are a cornerstone of experimental design on the web, with wide-ranging applications and use-cases. The statistical $t$-test comparing differences in means is the most commonly used method for assessing treatment effects, often…

统计方法学 · 统计学 2025-02-25 Olivier Jeunen

In an A/B test, the typical objective is to measure the total average treatment effect (TATE), which measures the difference between the average outcome if all users were treated and the average outcome if all users were untreated. However,…

应用统计 · 统计学 2020-04-28 David Holtz , Sinan Aral

A challenge that machine learning practitioners in the industry face is the task of selecting the best model to deploy in production. As a model is often an intermediate component of a production system, online controlled experiments such…

A/B testing, a widely used form of Randomized Controlled Trial (RCT), is a fundamental tool in business data analysis and experimental design. However, despite its intent to maintain randomness, A/B testing often faces challenges that…

统计方法学 · 统计学 2024-08-13 Zihao Zheng , Carol Liu

A/B experiments are commonly used in research to compare the effects of changing one or more variables in two different experimental groups - a control group and a treatment group. While the benefits of using A/B experiments are widely…

软件工程 · 计算机科学 2023-09-26 Andrew Hornback , Sungeun An , Scott Bunin , Stephen Buckley , John Kos , Ashok Goel

A critical challenge in recommender systems is to establish reliable relationships between offline and online metrics that predict real-world performance. Motivated by recent advances in Pareto front approximation, we introduce a pragmatic…

信息检索 · 计算机科学 2025-07-15 Timo Wilm , Philipp Normann