中文
相关论文

相关论文: Powerful A/B-Testing Metrics and Where to Find The…

200 篇论文

Offline evaluation plays a central role in benchmarking recommender systems when online testing is impractical or risky. However, it is susceptible to two key sources of bias: exposure bias, where users only interact with items they are…

信息检索 · 计算机科学 2025-08-12 Bruno L. Pereira , Alan Said , Rodrygo L. T. Santos

AB testing aids business operators with their decision making, and is considered the gold standard method for learning from data to improve digital user experiences. However, there is usually a gap between the requirements of practitioners,…

Making ideal decisions as a product leader in a web-facing company is extremely difficult. In addition to navigating the ambiguity of customer satisfaction and achieving business goals, one must also pave a path forward for ones' products…

Online controlled experiments (a.k.a. A/B testing) have been used as the mantra for data-driven decision making on feature changing and product shipping in many Internet companies. However, it is still a great challenge to systematically…

应用统计 · 统计学 2018-08-16 Yuxiang Xie , Nanyu Chen , Xiaolin Shi

Online advertisements have become one of today's most widely used tools for enhancing businesses partly because of their compatibility with A/B testing. A/B testing allows sellers to find effective advertisement strategies such as ad…

机器学习 · 计算机科学 2020-10-22 Akira Matsui , Daisuke Moriwaki

Utilizing randomized experiments to evaluate the effect of short-term treatments on the short-term outcomes has been well understood and become the golden standard in industrial practice. However, as service systems become increasingly…

统计方法学 · 统计学 2025-09-10 Shuze Chen , David Simchi-Levi , Chonghuan Wang

A/B testing is an important decision-making tool in product development for evaluating user engagement or satisfaction from a new service, feature or product. The goal of A/B testing is to estimate the average treatment effects (ATE) of a…

统计方法学 · 统计学 2020-08-21 Yifan Zhou , Yang Liu , Ping Li , Feifang Hu

The effectiveness of recommendation systems is pivotal to user engagement and satisfaction in online platforms. As these recommendation systems increasingly influence user choices, their evaluation transcends mere technical performance and…

信息检索 · 计算机科学 2024-01-15 Aryan Jadon , Avinash Patil

Online experiments %in which experimental units receive a sequence of treatments over time are frequently employed in many technological companies to evaluate the performance of a newly developed policy, product, or treatment relative to a…

计量经济学 · 经济学 2025-01-14 Ke Sun , Linglong Kong , Hongtu Zhu , Chengchun Shi

AB-testing is a very popular technique in web companies since it makes it possible to accurately predict the impact of a modification with the simplicity of a random split across users. One of the critical aspects of an AB-test is its…

机器学习 · 统计学 2015-02-02 Cyrille Dubarry

Participants in online experiments often enroll over time, which can compromise sample representativeness due to temporal shifts in covariates. This issue is particularly critical in A/B tests, online controlled experiments extensively used…

综合经济学 · 经济学 2026-03-30 Chen Wang , Shichao Han , Shan Huang

As the demand for mobile connectivity continues to grow, there is a strong need to evaluate the performance of Mobile Broadband (MBB) networks. In the last years, mobile "speed", quantified most commonly by data rate, gained popularity as…

网络与互联网体系结构 · 计算机科学 2018-01-31 Cise Midoglu , Leonhard Wimmer , Andra Lutu , Ozgu Alay , Carsten Griwodz

Effective exploration is believed to positively influence the long-term user experience on recommendation platforms. Determining its exact benefits, however, has been challenging. Regular A/B tests on exploration often measure neutral or…

We describe how to calculate standard errors for A/B tests that include clustered data, ratio metrics, and/or covariate adjustment. We may do this for power analysis/sample size calculations prior to running an experiment using historical…

统计方法学 · 统计学 2024-06-12 Tim Hesterberg , Ben Knight

Many organizations utilize large-scale online controlled experiments (OCEs) to accelerate innovation. Having high statistical power to detect small differences between control and treatment accurately is critical, as even small changes in…

应用统计 · 统计学 2020-09-11 Ali Mahmoudzadeh , Sophia Liu , Sol Sadeghi , Paul Luo Li , Somit Gupta

A/B tests, also known as randomized controlled experiments (RCTs), are the gold standard for evaluating the impact of new policies, products, or decisions. However, these tests can be costly in terms of time and resources, potentially…

机器学习 · 统计学 2025-01-03 Shima Nassiri , Mohsen Bayati , Joe Cooprider

News recommenders help users to find relevant online content and have the potential to fulfill a crucial role in a democratic society, directing the scarce attention of citizens towards the information that is most important to them.…

信息检索 · 计算机科学 2020-12-21 Sanne Vrijenhoek , Mesut Kaya , Nadia Metoui , Judith Möller , Daan Odijk , Natali Helberger

A/B testing is an important decision making tool in product development because can provide an accurate estimate of the average treatment effect of a new features, which allows developers to understand how the business impact of new changes…

应用统计 · 统计学 2019-03-22 Guillaume Saint-Jacques , James Eric Sorenson , Nanyu Chen , Ya Xu

In many industry settings, online controlled experimentation (A/B test) has been broadly adopted as the gold standard to measure product or feature impacts. Most research has primarily focused on user engagement type metrics, specifically…

统计方法学 · 统计学 2020-10-30 Weinan Wang , Xi Zhang

Marketers often use A/B testing as a tool to compare marketing treatments in a test stage and then deploy the better-performing treatment to the remainder of the consumer population. While these tests have traditionally been analyzed using…

应用统计 · 统计学 2020-12-03 Elea McDonnell Feit , Ron Berman