中文
相关论文

相关论文: Quantifying the Value of Iterative Experimentation

200 篇论文

Experimentation in online digital platforms is used to inform decision making. Specifically, the goal of many experiments is to optimize a metric of interest. Null hypothesis statistical testing can be ill-suited to this task, as it is…

统计方法学 · 统计学 2024-12-10 Timothy Sudijono , Simon Ejdemyr , Apoorva Lal , Martin Tingley

With the extensive use of digital devices, online experimental platforms are commonly used to conduct experiments to collect data for evaluating different variations of products, algorithms, and interface designs, a.k.a., A/B tests. In…

统计方法学 · 统计学 2024-07-09 Qiong Zhang , Lulu Kang , Xinwei Deng

Online controlled experimentation is widely adopted for evaluating new features in the rapid development cycle for web products and mobile applications. Measurement of the overall experiment sample is a common practice to quantify the…

人机交互 · 计算机科学 2022-01-27 Zhenyu Zhao , Yan He , Miao Chen

A/B tests serve the purpose of reliably identifying the effect of changes introduced in online services. It is common for online platforms to run a large number of simultaneous experiments by splitting incoming user traffic randomly in…

We have seen a massive growth of online experiments at LinkedIn, and in industry at large. It is now more important than ever to create an intelligent A/B platform that can truly democratize A/B testing by allowing everyone to make quality…

应用统计 · 统计学 2018-08-02 Nanyu Chen , Min Liu , Ya Xu

Technology firms conduct randomized controlled experiments ("A/B tests") to learn which actions to take to improve business outcomes. In firms with mature experimentation platforms, experimentation programs can consist of many thousands of…

统计方法学 · 统计学 2025-05-30 Winston Chou , Colin Gray , Nathan Kallus , Aurélien Bibaut , Simon Ejdemyr

Online experiments are the gold standard for evaluating impact on user experience and accelerating innovation in software. However, since experiments are typically limited in duration, observed treatment effects are not always permanently…

人机交互 · 计算机科学 2021-02-26 Soheil Sadeghi , Somit Gupta , Stefan Gramatovici , Jiannan Lu , Hao Ai , Ruhan Zhang

Estimating the effects of long-term treatments through A/B testing is challenging. Treatments, such as updates to product functionalities, user interface designs, and recommendation algorithms, are intended to persist within the system for…

计量经济学 · 经济学 2025-12-30 Shan Huang , Chen Wang , Yuan Yuan , Jinglong Zhao , Brocco , Zhang

Development of the majority of the leading web services and software products today is generally guided by data-driven decisions based on evaluation that ensures a steady stream of updates, both in terms of quality and quantity. Large…

人机交互 · 计算机科学 2018-09-05 Roman Budylin , Alexey Drutsa , Gleb Gusev , Pavel Serdyukov , Igor Yashkov

A/B testing experiment is a widely adopted method for evaluating UI/UX design decisions in modern web applications. Yet, traditional A/B testing remains constrained by its dependence on the large-scale and live traffic of human…

In recent years, a vivid interest in hybrid development methods has been observed as practitioners combine various approaches to software creation to improve productivity, product quality, and adaptability of the process to react to change.…

软件工程 · 计算机科学 2021-03-08 Rafał Włodarski , Jean-Rémy Falleri , Corinne Parvéry

In industry, online randomized controlled experiment (a.k.a. A/B experiment) is a standard approach to measure the impact of a causal change. These experiments have small treatment effect to reduce the potential blast radius. As a result,…

计量经济学 · 经济学 2025-05-29 Tanmoy Das , Dohyeon Lee , Arnab Sinha

When developing a new networking algorithm, it is established practice to run a randomized experiment, or A/B test, to evaluate its performance. In an A/B test, traffic is randomly allocated between a treatment group, which uses the new…

网络与互联网体系结构 · 计算机科学 2021-10-04 Bruce Spang , Veronica Hannan , Shravya Kunamalla , Te-Yuan Huang , Nick McKeown , Ramesh Johari

Experimental testing is vital in the optimization of web applications, and as such A/B testing has been widely adopted as a methodology for determining optimal content for many web applications. While some testing platforms provide…

统计方法学 · 统计学 2017-10-04 Ian E. Fellows

Autonomous agents powered by large language models (LLMs) show significant potential for achieving high autonomy in various scenarios such as software development. Recent research has shown that LLM agents can leverage past experiences to…

计算与语言 · 计算机科学 2024-05-08 Chen Qian , Jiahao Li , Yufan Dang , Wei Liu , YiFei Wang , Zihao Xie , Weize Chen , Cheng Yang , Yingli Zhang , Zhiyuan Liu , Maosong Sun

Controlled experiments (A/B tests or randomized field experiments) are the de facto standard to make data-driven decisions when implementing changes and observing customer responses. The methodology to analyze such experiments should be…

应用统计 · 统计学 2020-03-06 Shafi Kamalbasha , Manuel J. A. Eugster

Evaluation plays a crucial role in the development of ranking algorithms on search and recommender systems. It enables online platforms to create user-friendly features that drive commercial success in a steady and effective manner. The…

信息检索 · 计算机科学 2025-08-04 Qing Zhang , Alex Deng , Michelle Du , Huiji Gao , Liwei He , Sanjeev Katariya

Large language models (LLMs) are now used in multi-turn workflows, but we still lack a clear way to measure when iteration helps and when it hurts. We present an evaluation framework for iterative refinement that spans ideation, code, and…

人工智能 · 计算机科学 2025-09-16 Shashidhar Reddy Javaji , Bhavul Gauri , Zining Zhu

Online controlled experiments, such as A/B-tests, are commonly used by modern tech companies to enable continuous system improvements. Despite their paramount importance, A/B-tests are expensive: by their very definition, a percentage of…

机器学习 · 计算机科学 2024-01-09 Shubham Baweja , Neeti Pokharna , Aleksei Ustimenko , Olivier Jeunen

The widespread adoption of online randomized controlled experiments (A/B Tests) for decision-making has created ongoing capacity constraints which necessitate interim analyses. As a consequence, platform users are increasingly motivated to…