中文
相关论文

相关论文: Evaluating Decision Rules Across Many Weak Experim…

200 篇论文

The goal of this article is to investigate how human participants allocate their limited time to decisions with different properties. We report the results of two behavioral experiments. In each trial of the experiments, the participant…

神经元与认知 · 定量生物学 2016-07-20 Arash Khodadadi , Pegah Fakhari , Jerome R. Busemeyer

Heavy-tailed metrics are common and often critical to product evaluation in the online world. While we may have samples large enough for Central Limit Theorem to kick in, experimentation is challenging due to the wide confidence interval of…

应用统计 · 统计学 2019-05-23 Jason , Wang , Pauline Burke

E-commerce companies have a number of online products, such as organic search, sponsored search, and recommendation modules, to fulfill customer needs. Although each of these products provides a unique opportunity for users to interact with…

应用统计 · 统计学 2020-06-23 Xuan Yin , Liangjie Hong

AB testing aids business operators with their decision making, and is considered the gold standard method for learning from data to improve digital user experiences. However, there is usually a gap between the requirements of practitioners,…

It is widely believed that one's peers influence product adoption behaviors. This relationship has been linked to the number of signals a decision-maker receives in a social network. But it is unclear if these same principles hold when the…

社会与信息网络 · 计算机科学 2020-09-09 Soumajyoti Sarkar , Ashkan Aleali , Paulo Shakarian , Mika Armenta , Danielle Sanchez , Kiran Lakkaraju

Failure to accurately measure the outcomes of an experiment can lead to bias and incorrect conclusions. Online controlled experiments (aka AB tests) are increasingly being used to make decisions to improve websites as well as mobile and…

其他计算机科学 · 计算机科学 2019-04-01 Jayant Gupchup , Yasaman Hosseinkashi , Pavel Dmitriev , Daniel Schneider , Ross Cutler , Andrei Jefremov , Martin Ellis

Top-N item recommendation has been a widely studied task from implicit feedback. Although much progress has been made with neural methods, there is increasing concern on appropriate evaluation of recommendation algorithms. In this paper, we…

信息检索 · 计算机科学 2020-10-12 Wayne Xin Zhao , Junhua Chen , Pengfei Wang , Qi Gu , Ji-Rong Wen

Effective exploration is believed to positively influence the long-term user experience on recommendation platforms. Determining its exact benefits, however, has been challenging. Regular A/B tests on exploration often measure neutral or…

Adaptive experiments are used extensively in online platforms, healthcare and biotechnology, and a variety of other settings. In many of these applications, the main goal is not to precisely estimate a treatment effect, but to demonstrate…

Current practice for evaluating recommender systems typically focuses on point estimates of user-oriented effectiveness metrics or business metrics, sometimes combined with additional metrics for considerations such as diversity and…

信息检索 · 计算机科学 2023-09-13 Michael D. Ekstrand , Ben Carterette , Fernando Diaz

Recent research in causal inference under network interference has explored various experimental designs and estimation techniques to address this issue. However, existing methods, which typically rely on single experiments, often reach a…

统计方法学 · 统计学 2025-03-10 Qianyi Chen , Bo Li

A B testing serves as the gold standard for large scale, data driven decision making in online businesses. To mitigate metric variability and enhance testing sensitivity, control variates and regression adjustment have emerged as prominent…

统计方法学 · 统计学 2025-10-13 Yu Zhang , Bokui Wan , Yongli Qin

A/B testing is an important decision-making tool in product development for evaluating user engagement or satisfaction from a new service, feature or product. The goal of A/B testing is to estimate the average treatment effects (ATE) of a…

统计方法学 · 统计学 2020-08-21 Yifan Zhou , Yang Liu , Ping Li , Feifang Hu

Numerical evaluations with comparisons to baselines play a central role when judging research in recommender systems. In this paper, we show that running baselines properly is difficult. We demonstrate this issue on two extensively studied…

信息检索 · 计算机科学 2019-05-07 Steffen Rendle , Li Zhang , Yehuda Koren

Inspired by the legacy of the Netflix contest, we provide an overview of what has been learned---from our own efforts, and those of others---concerning the problems of collaborative filtering and recommender systems. The data set consists…

统计方法学 · 统计学 2012-07-25 Andrey Feuerverger , Yu He , Shashi Khatri

Experimentation is widely utilized for causal inference and data-driven decision-making across disciplines. In an A/B experiment, for example, an online business randomizes two different treatments (e.g., website designs) to their customers…

统计方法学 · 统计学 2025-01-15 Wenxuan Guo , JungHo Lee , Panos Toulis

This paper studies prototypical strategies to sequentially aggregate independent decisions. We consider a collection of agents, each performing binary hypothesis testing and each obtaining a decision over time. We assume the agents are…

应用统计 · 统计学 2010-08-30 Sandra H. Dandach , Ruggero Carli , Francesco Bullo

A/B testing has become the gold standard for policy evaluation in modern technological industries. Motivated by the widespread use of switchback experiments in A/B testing, this paper conducts a comprehensive comparative analysis of various…

机器学习 · 统计学 2025-08-29 Qianglin Wen , Chengchun Shi , Ying Yang , Niansheng Tang , Hongtu Zhu

Online experiments such as Randomised Controlled Trials (RCTs) or A/B-tests are the bread and butter of modern platforms on the web. They are conducted continuously to allow platforms to estimate the causal effect of replacing system…

机器学习 · 计算机科学 2023-04-24 Olivier Jeunen

In online randomized experiments or A/B tests, accurate predictions of participant inclusion rates are of paramount importance. These predictions not only guide experimenters in optimizing the experiment's duration but also enhance the…

统计方法学 · 统计学 2024-02-06 Lorenzo Masoero , Mario Beraha , Thomas Richardson , Stefano Favaro