中文
相关论文

相关论文: Quantifying the Value of Iterative Experimentation

200 篇论文

This study presents a theoretical analysis on the efficiency of interleaving, an efficient online evaluation method for rankings. Although interleaving has already been applied to production systems, the source of its high efficiency has…

信息检索 · 计算机科学 2023-06-21 Kojiro Iizuka , Hajime Morita , Makoto P. Kato

In this paper, we examine the biases that arise when firms run A/B tests on continuous parameters to estimate global treatment effects on performance metrics of interest; we particularly focus on price experiments to measure the price…

统计方法学 · 统计学 2026-01-22 Ramesh Johari , Orrie B. Page , Gabriel Y. Weintraub

With the growing needs of online A/B testing to support the innovation in industry, the opportunity cost of running an experiment becomes non-negligible. Therefore, there is an increasing demand for an efficient continuous monitoring…

机器学习 · 计算机科学 2023-04-04 Runzhe Wan , Yu Liu , James McQueen , Doug Hains , Rui Song

Motivated by the widespread adoption of iterative project management techniques, we study the effects of workflow -- iterative or sequential -- on innovative behavior and performance. We conduct a series of laboratory experiments. Our first…

综合经济学 · 经济学 2026-03-03 Evgeny Kagan , Christian Jost , Tobias Lieberum , Sebastian Schiffels

A/B testing is ubiquitous within the machine learning and data science operations of internet companies. Generically, the idea is to perform a statistical test of the hypothesis that a new feature is better than the existing platform---for…

统计理论 · 数学 2017-10-11 David Goldberg , James E. Johndrow

A/B testing plays a central role in data-driven product development, guiding launch decisions for new features and designs. However, treatment effect estimates are often noisy due to short horizons, early stopping, and slowly accumulating…

统计方法学 · 统计学 2025-11-27 Xinran Li

Machine learning workflow development is anecdotally regarded to be an iterative process of trial-and-error with humans-in-the-loop. However, we are not aware of quantitative evidence corroborating this popular belief. A quantitative…

机器学习 · 计算机科学 2018-05-21 Doris Xin , Litian Ma , Shuchen Song , Aditya Parameswaran

A/B testing is critical for modern technological companies to evaluate the effectiveness of newly developed products against standard baselines. This paper studies optimal designs that aim to maximize the amount of information obtained from…

统计方法学 · 统计学 2023-11-07 Ting Li , Chengchun Shi , Jianing Wang , Fan Zhou , Hongtu Zhu

Effective exploration is believed to positively influence the long-term user experience on recommendation platforms. Determining its exact benefits, however, has been challenging. Regular A/B tests on exploration often measure neutral or…

Online evaluation of machine learning models is typically conducted through A/B experiments. Sequential statistical tests are valuable tools for analysing these experiments, as they enable researchers to stop data collection early without…

统计方法学 · 统计学 2025-10-08 Alexey Kurennoy , Majed Dodin , Tural Gurbanov , Ana Peleteiro Ramallo

Utilizing randomized experiments to evaluate the effect of short-term treatments on the short-term outcomes has been well understood and become the golden standard in industrial practice. However, as service systems become increasingly…

统计方法学 · 统计学 2025-09-10 Shuze Chen , David Simchi-Levi , Chonghuan Wang

The rollout of new versions of a feature in modern applications is a manual multi-stage process, as the feature is released to ever larger groups of users, while its performance is carefully monitored. This kind of A/B testing is…

机器学习 · 计算机科学 2018-05-29 Andrés Muñoz Medina , Sergei Vassilvitskii , Dong Yin

Software companies have widely used online A/B testing to evaluate the impact of a new technology by offering it to groups of users and comparing it against the unmodified product. However, running online A/B testing needs not only efforts…

软件工程 · 计算机科学 2024-08-12 Jie JW Wu

Experimentation is widely utilized for causal inference and data-driven decision-making across disciplines. In an A/B experiment, for example, an online business randomizes two different treatments (e.g., website designs) to their customers…

统计方法学 · 统计学 2025-01-15 Wenxuan Guo , JungHo Lee , Panos Toulis

AB testing aids business operators with their decision making, and is considered the gold standard method for learning from data to improve digital user experiences. However, there is usually a gap between the requirements of practitioners,…

The importance of Facebook advertising has risen dramatically in recent years, with the platform accounting for almost 20% of the global online ad spend in 2017. An important consideration in advertising is incrementality: how much of the…

统计方法学 · 统计学 2018-07-12 C. H. Bryan Liu , Elaine M. Bettaney , Benjamin Paul Chamberlain

Smartphone manufacturers continue to release new models annually, yet the pace of meaningful innovation has slowed, with most changes limited to incremental updates in design, performance, or software. This study examines whether such…

新兴技术 · 计算机科学 2026-03-04 Chandima Wickramatunga , Ruwan Nagahawatta , Anagi Gamachchi , Chintha Kaluarachchi

Test-Driven Development (TDD) has been claimed to increase external software quality. However, the extent to which TDD increases external quality has been seldom studied in industrial experiments. We conduct four industrial experiments in…

软件工程 · 计算机科学 2018-07-23 Adrian Santos , Janne Jarvinen , Jari Partanen , Markku Oivo , Natalia Juristo

A/B testing is the foundation of decision-making in online platforms, yet social products often suffer from network interference: user interactions cause treatment effects to spill over into the control group. Such spillovers bias causal…

社会与信息网络 · 计算机科学 2026-02-10 Xu Min , Zhaoxu Yang , Kaixuan Tan , Juan Yan , Xunbin Xiong , Zihao Zhu , Kaiyu Zhu , Fenglin Cui , Yang Yang , Sihua Yang , Jianhui Bu

Online controlled experiments, or A/B tests, are large-scale randomized trials in digital environments. This paper investigates the estimands of the difference-in-means estimator in these experiments, focusing on scenarios with repeated…

统计方法学 · 统计学 2024-11-12 Sebastian Ankargren , Mattias Frånberg , Mårten Schultzberg