English
Related papers

Related papers: Powerful A/B-Testing Metrics and Where to Find The…

200 papers

Current practice for evaluating recommender systems typically focuses on point estimates of user-oriented effectiveness metrics or business metrics, sometimes combined with additional metrics for considerations such as diversity and…

Information Retrieval · Computer Science 2023-09-13 Michael D. Ekstrand , Ben Carterette , Fernando Diaz

Our opinions, which things we like or dislike, depend on the opinions of those around us. Nowadays, we are influenced by the opinions of online strangers, expressed in comments and ratings on online platforms. Here, we perform novel…

Innovations across science and industry are evaluated using randomized trials (a.k.a. A/B tests). While simple and robust, such static designs are inefficient or infeasible for testing many hypotheses. Adaptive designs can greatly improve…

Machine Learning · Computer Science 2024-08-09 Jimmy Wang , Ethan Che , Daniel R. Jiang , Hongseok Namkoong

To identify the most appropriate recommendation model for an e-commerce business, a live evaluation should be performed on the shopping website to measure the influence of personalization in real-time. The aim of this paper is to introduce…

Information Retrieval · Computer Science 2019-01-28 Namrata Chaudhary , Drimik Roy Chowdhury

Conformance checking techniques aim to collate observed process behavior with normative/modeled process models. The majority of existing approaches focuses on completed process executions, i.e., offline conformance checking. Recently, novel…

Logic in Computer Science · Computer Science 2022-11-23 Daniel Schuster , Gero J. Kolhof

In online multiple testing, an a priori unknown number of hypotheses are tested sequentially, i.e. at each time point a test decision for the current hypothesis has to be made using only the data available so far. Although many powerful…

Methodology · Statistics 2025-03-11 Vincent Jankovic , Lasse Fischer , Werner Brannath

Recommender systems are crucial tools to overcome the information overload brought about by the Internet. Rigorous tests are needed to establish to what extent sophisticated methods can improve the quality of the predictions. Here we…

Information Retrieval · Computer Science 2007-09-19 Marcel Blattner , Alexander Hunziker , Paolo Laureti

Randomized experiments, or A/B testing, are the gold standard for evaluating interventions, yet they remain underutilized in inventory management. This study addresses this gap by analyzing A/B testing strategies in multi-item, multi-period…

Methodology · Statistics 2026-02-03 Xinqi Chen , Xingyu Bai , Zeyu Zheng , Nian Si

Online controlled experiments (also known as A/B Testing) have been viewed as a golden standard for large data-driven companies since the last few decades. The most common A/B testing framework adopted by many companies use "average…

Methodology · Statistics 2021-10-15 Yihan Bao , Shichao Han , Yong Wang

Traditional artificial-star tests are widely applied to photometry in crowded stellar fields. However, to obtain reliable binary fractions (and their uncertainties) of remote, dense, and rich star clusters, one needs to recover huge numbers…

Instrumentation and Methods for Astrophysics · Physics 2015-05-20 Yi Hu , Licai Deng , Richard de Grijs , Qiang Liu

Standard A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inferring experiment statistics such as the average…

Machine Learning · Computer Science 2025-06-04 Qining Zhang , Tanner Fiez , Yi Liu , Wenyang Liu

Classification systems are evaluated in a countless number of papers. However, we find that evaluation practice is often nebulous. Frequently, metrics are selected without arguments, and blurry terminology invites misconceptions. For…

Machine Learning · Computer Science 2024-07-03 Juri Opitz

A/B tests are often required to be conducted on subjects that might have social connections. For e.g., experiments on social media, or medical and social interventions to control the spread of an epidemic. In such settings, the SUTVA…

Machine Learning · Computer Science 2024-04-17 Shiv Shankar , Ritwik Sinha , Yash Chandak , Saayan Mitra , Madalina Fiterau

Millions of mobile apps are available in app stores, such as Apple's App Store and Google Play. For a mobile app, it would be increasingly challenging to stand out from the enormous competitors and become prevalent among users. Good user…

Software Engineering · Computer Science 2020-08-25 Cuiyun Gao , Jichuan Zeng , Zhiyuan Wen , David Lo , Xin Xia , Irwin King , Michael R. Lyu

This paper proposes new nonparametric diagnostic tools to assess the asymptotic validity of different treatment effects estimators that rely on the correct specification of the propensity score. We derive a particular restriction relating…

Methodology · Statistics 2019-02-11 Pedro H. C. Sant'Anna , Xiaojun Song

Machine learning models are being used extensively in many important areas, but there is no guarantee a model will always perform well or as its developers intended. Understanding the correctness of a model is crucial to prevent potential…

Machine Learning · Computer Science 2021-04-13 Huong Ha , Sunil Gupta , Santu Rana , Svetha Venkatesh

In the past two decades, AB testing has proliferated to optimise products in digital domains. Traditional AB tests use fixed-horizon testing, determining the sample size of the experiment and continuing until the experiment has concluded.…

Methodology · Statistics 2023-11-01 Daniel Beasley

The aim of online monitoring is to issue an alarm as soon as there is significant evidence in the collected observations to suggest that the underlying data generating mechanism has changed. This work is concerned with open-end,…

Statistics Theory · Mathematics 2020-07-21 Mark Holmes , Ivan Kojadinovic

With the growing needs of online A/B testing to support the innovation in industry, the opportunity cost of running an experiment becomes non-negligible. Therefore, there is an increasing demand for an efficient continuous monitoring…

Machine Learning · Computer Science 2023-04-04 Runzhe Wan , Yu Liu , James McQueen , Doug Hains , Rui Song

There are a large number of competing ADXs on the Internet. It is the primary demand to identify and compare the advertising performance of ADX. Traditional method relies on training artificial online personas to represent behavioral…

Computers and Society · Computer Science 2018-03-19 Nanxi Huang , Chunxi Li , Yongxiang Zhao , Yuchun Guo
‹ Prev 1 8 9 10 Next ›