中文
相关论文

相关论文: Evaluating Decision Rules Across Many Weak Experim…

200 篇论文

Making a decision is often a matter of listing and comparing positive and negative arguments. In such cases, the evaluation scale for decisions should be considered bipolar, that is, negative and positive values should be explicitly…

人工智能 · 计算机科学 2014-01-16 Didier Dubois , Hélène Fargier , Jean-François Bonnefon

Evaluating the causal effect of recommendations is an important objective because the causal effect on user interactions can directly leads to an increase in sales and user engagement. To select an optimal recommendation model, it is common…

机器学习 · 计算机科学 2021-07-16 Masahiro Sato

We describe how to calculate standard errors for A/B tests that include clustered data, ratio metrics, and/or covariate adjustment. We may do this for power analysis/sample size calculations prior to running an experiment using historical…

统计方法学 · 统计学 2024-06-12 Tim Hesterberg , Ben Knight

Online controlled experiments, now commonly known as A/B testing, are crucial to causal inference and data driven decision making in many internet based businesses. While a simple comparison between a treatment (the feature under test) and…

应用统计 · 统计学 2015-01-05 Yu Guo , Alex Deng

This paper introduces a method for linking technological improvement rates (i.e. Moore's Law) and technology adoption curves (i.e. S-Curves). There has been considerable research surrounding Moore's Law and the generalized versions applied…

计量经济学 · 经济学 2018-05-17 Christopher L. Benson , Christopher L. Magee

A/B testing is a core tool for decision-making in business experimentation, particularly in digital platforms and marketplaces. Practitioners often prioritize lift in performance metrics while seeking to control the costs of false…

统计方法学 · 统计学 2025-08-21 Pallavi Basu , Ron Berman

Online experimentation, also known as A/B testing, is the gold standard for measuring product impacts and making business decisions in the tech industry. The validity and utility of experiments, however, hinge on unbiasedness and sufficient…

应用统计 · 统计学 2020-12-17 Min Liu , Jialiang Mao , Kang Kang

A/B tests are randomized experiments frequently used by companies that offer services on the Web for assessing the impact of new features. During an experiment, each user is randomly redirected to one of two versions of the website, called…

社会与信息网络 · 计算机科学 2021-08-12 Francisco Galuppo Azevedo , Bruno Demattos Nogueira , Fabricio Murai , Ana Paula Couto da Silva

Randomized experimentation (also known as A/B testing or bucket testing) is widely used in the internet industry to measure the metric impact obtained by different treatment variants. A/B tests identify the treatment variant showing the…

统计方法学 · 统计学 2020-12-23 Ye Tu , Kinjal Basu , Cyrus DiCiccio , Romil Bansal , Preetam Nandy , Padmini Jaikumar , Shaunak Chatterjee

In this paper, we propose an approach to analyze the performance and the added value of automatic recommender systems in an industrial context. We show that recommender systems are multifaceted and can be organized around 4 structuring…

信息检索 · 计算机科学 2015-03-13 Frank Meyer , Françoise Fessant , Fabrice Clérot , Eric Gaussier

In this paper, we examine the statistical soundness of comparative assessments within the field of recommender systems in terms of reliability and human uncertainty. From a controlled experiment, we get the insight that users provide…

人机交互 · 计算机科学 2017-06-28 Kevin Jasberg , Sergej Sizov

We study ratio metrics in A/B testing at the presence of correlation among observations coming from the same user and provides practical guidance especially when two metrics contradict each other. We propose new estimating methods to…

应用统计 · 统计学 2020-07-24 Keyu Nie , Yinfei Kong , Ted Tao Yuan , Pauline Berry Burke

Evaluation plays a crucial role in the development of ranking algorithms on search and recommender systems. It enables online platforms to create user-friendly features that drive commercial success in a steady and effective manner. The…

信息检索 · 计算机科学 2025-08-04 Qing Zhang , Alex Deng , Michelle Du , Huiji Gao , Liwei He , Sanjeev Katariya

Digital firms routinely run many online experiments on shared user populations. When product decisions are compositional, such as combinations of interface elements, flows, messages, or incentives, the number of feasible interventions grows…

机器学习 · 统计学 2026-04-13 Xin Wen , Xi Chen , Will Wei Sun , Yichen Zhang

Experimental testing is vital in the optimization of web applications, and as such A/B testing has been widely adopted as a methodology for determining optimal content for many web applications. While some testing platforms provide…

统计方法学 · 统计学 2017-10-04 Ian E. Fellows

Online controlled experiments, also known as A/B testing, are the digital equivalent of randomized controlled trials for estimating the impact of marketing campaigns on website visitors. Stratified sampling is a traditional technique for…

In this paper, a likelihood based evidence acquisition approach is proposed to acquire evidence from experts'assessments as recorded in historical datasets. Then a data-driven evidential reasoning rule based model is introduced to R&D…

数字图书馆 · 计算机科学 2018-11-21 Fang Liu , Yu-wang Chen , Jian-bo Yang , Dong-ling Xu , Wei-shu Liu

Evaluating retrieval-ranking systems is crucial for developing high-performing models. While online A/B testing is the gold standard, its high cost and risks to user experience require effective offline methods. However, relying on…

信息检索 · 计算机科学 2025-04-08 Seyedeh Baharan Khatami , Sayan Chakraborty , Ruomeng Xu , Babak Salimi

A/B testing methodology is generally performed by private companies to increase user engagement and satisfaction about online features. Their usage is far from being transparent and may undermine user autonomy (e.g. polarizing individual…

社会与信息网络 · 计算机科学 2024-05-03 Matteo Ottaviani , Stefan M. Herzog , Pietro Leonardo Nickl , Philipp Lorenz-Spreen

In an A/B test, the typical objective is to measure the total average treatment effect (TATE), which measures the difference between the average outcome if all users were treated and the average outcome if all users were untreated. However,…

应用统计 · 统计学 2020-04-28 David Holtz , Sinan Aral