English
Related papers

Related papers: Evaluating Decision Rules Across Many Weak Experim…

200 papers

Making a decision is often a matter of listing and comparing positive and negative arguments. In such cases, the evaluation scale for decisions should be considered bipolar, that is, negative and positive values should be explicitly…

Artificial Intelligence · Computer Science 2014-01-16 Didier Dubois , Hélène Fargier , Jean-François Bonnefon

Evaluating the causal effect of recommendations is an important objective because the causal effect on user interactions can directly leads to an increase in sales and user engagement. To select an optimal recommendation model, it is common…

Machine Learning · Computer Science 2021-07-16 Masahiro Sato

We describe how to calculate standard errors for A/B tests that include clustered data, ratio metrics, and/or covariate adjustment. We may do this for power analysis/sample size calculations prior to running an experiment using historical…

Methodology · Statistics 2024-06-12 Tim Hesterberg , Ben Knight

Online controlled experiments, now commonly known as A/B testing, are crucial to causal inference and data driven decision making in many internet based businesses. While a simple comparison between a treatment (the feature under test) and…

Applications · Statistics 2015-01-05 Yu Guo , Alex Deng

This paper introduces a method for linking technological improvement rates (i.e. Moore's Law) and technology adoption curves (i.e. S-Curves). There has been considerable research surrounding Moore's Law and the generalized versions applied…

Econometrics · Economics 2018-05-17 Christopher L. Benson , Christopher L. Magee

A/B testing is a core tool for decision-making in business experimentation, particularly in digital platforms and marketplaces. Practitioners often prioritize lift in performance metrics while seeking to control the costs of false…

Methodology · Statistics 2025-08-21 Pallavi Basu , Ron Berman

Online experimentation, also known as A/B testing, is the gold standard for measuring product impacts and making business decisions in the tech industry. The validity and utility of experiments, however, hinge on unbiasedness and sufficient…

Applications · Statistics 2020-12-17 Min Liu , Jialiang Mao , Kang Kang

A/B tests are randomized experiments frequently used by companies that offer services on the Web for assessing the impact of new features. During an experiment, each user is randomly redirected to one of two versions of the website, called…

Social and Information Networks · Computer Science 2021-08-12 Francisco Galuppo Azevedo , Bruno Demattos Nogueira , Fabricio Murai , Ana Paula Couto da Silva

Randomized experimentation (also known as A/B testing or bucket testing) is widely used in the internet industry to measure the metric impact obtained by different treatment variants. A/B tests identify the treatment variant showing the…

In this paper, we propose an approach to analyze the performance and the added value of automatic recommender systems in an industrial context. We show that recommender systems are multifaceted and can be organized around 4 structuring…

Information Retrieval · Computer Science 2015-03-13 Frank Meyer , Françoise Fessant , Fabrice Clérot , Eric Gaussier

In this paper, we examine the statistical soundness of comparative assessments within the field of recommender systems in terms of reliability and human uncertainty. From a controlled experiment, we get the insight that users provide…

Human-Computer Interaction · Computer Science 2017-06-28 Kevin Jasberg , Sergej Sizov

We study ratio metrics in A/B testing at the presence of correlation among observations coming from the same user and provides practical guidance especially when two metrics contradict each other. We propose new estimating methods to…

Applications · Statistics 2020-07-24 Keyu Nie , Yinfei Kong , Ted Tao Yuan , Pauline Berry Burke

Evaluation plays a crucial role in the development of ranking algorithms on search and recommender systems. It enables online platforms to create user-friendly features that drive commercial success in a steady and effective manner. The…

Information Retrieval · Computer Science 2025-08-04 Qing Zhang , Alex Deng , Michelle Du , Huiji Gao , Liwei He , Sanjeev Katariya

Digital firms routinely run many online experiments on shared user populations. When product decisions are compositional, such as combinations of interface elements, flows, messages, or incentives, the number of feasible interventions grows…

Machine Learning · Statistics 2026-04-13 Xin Wen , Xi Chen , Will Wei Sun , Yichen Zhang

Experimental testing is vital in the optimization of web applications, and as such A/B testing has been widely adopted as a methodology for determining optimal content for many web applications. While some testing platforms provide…

Methodology · Statistics 2017-10-04 Ian E. Fellows

Online controlled experiments, also known as A/B testing, are the digital equivalent of randomized controlled trials for estimating the impact of marketing campaigns on website visitors. Stratified sampling is a traditional technique for…

In this paper, a likelihood based evidence acquisition approach is proposed to acquire evidence from experts'assessments as recorded in historical datasets. Then a data-driven evidential reasoning rule based model is introduced to R&D…

Digital Libraries · Computer Science 2018-11-21 Fang Liu , Yu-wang Chen , Jian-bo Yang , Dong-ling Xu , Wei-shu Liu

Evaluating retrieval-ranking systems is crucial for developing high-performing models. While online A/B testing is the gold standard, its high cost and risks to user experience require effective offline methods. However, relying on…

Information Retrieval · Computer Science 2025-04-08 Seyedeh Baharan Khatami , Sayan Chakraborty , Ruomeng Xu , Babak Salimi

A/B testing methodology is generally performed by private companies to increase user engagement and satisfaction about online features. Their usage is far from being transparent and may undermine user autonomy (e.g. polarizing individual…

Social and Information Networks · Computer Science 2024-05-03 Matteo Ottaviani , Stefan M. Herzog , Pietro Leonardo Nickl , Philipp Lorenz-Spreen

In an A/B test, the typical objective is to measure the total average treatment effect (TATE), which measures the difference between the average outcome if all users were treated and the average outcome if all users were untreated. However,…

Applications · Statistics 2020-04-28 David Holtz , Sinan Aral