中文
相关论文

相关论文: Learning Metrics that Maximise Power for Accelerat…

200 篇论文

Experimental testing is vital in the optimization of web applications, and as such A/B testing has been widely adopted as a methodology for determining optimal content for many web applications. While some testing platforms provide…

统计方法学 · 统计学 2017-10-04 Ian E. Fellows

A/B testing is one of the most successful applications of statistical theory in modern Internet age. One problem of Null Hypothesis Statistical Testing (NHST), the backbone of A/B testing methodology, is that experimenters are not allowed…

应用统计 · 统计学 2016-02-18 Alex Deng , Jiannan Lu , Shouyuan Chen

We study ratio metrics in A/B testing at the presence of correlation among observations coming from the same user and provides practical guidance especially when two metrics contradict each other. We propose new estimating methods to…

应用统计 · 统计学 2020-07-24 Keyu Nie , Yinfei Kong , Ted Tao Yuan , Pauline Berry Burke

A/B testing is widely used in modern technology companies for policy evaluation and product deployment, with the goal of comparing the outcomes under a newly-developed policy against a standard control. Various causal inference and…

机器学习 · 统计学 2025-07-25 Jinjuan Wang , Qianglin Wen , Yu Zhang , Xiaodong Yan , Chengchun Shi

Experimentation platforms in industry must often deal with customer trust issues. Platforms must prove the validity of their claims as well as catch issues that arise. As a central quantity estimated by experimentation platforms, the…

统计方法学 · 统计学 2025-11-21 Kedar Karhadkar , Jack Klys , Daniel Ting , Artem Vorozhtsov , Houssam Nassif

A/B testing methodology is generally performed by private companies to increase user engagement and satisfaction about online features. Their usage is far from being transparent and may undermine user autonomy (e.g. polarizing individual…

社会与信息网络 · 计算机科学 2024-05-03 Matteo Ottaviani , Stefan M. Herzog , Pietro Leonardo Nickl , Philipp Lorenz-Spreen

In the absence of extensive human-annotated data for complex reasoning tasks, self-improvement -- where models are trained on their own outputs -- has emerged as a primary method for enhancing performance. However, the critical factors…

人工智能 · 计算机科学 2025-03-05 Weihao Zeng , Yuzhen Huang , Lulu Zhao , Yijun Wang , Zifei Shan , Junxian He

Utilizing randomized experiments to evaluate the effect of short-term treatments on the short-term outcomes has been well understood and become the golden standard in industrial practice. However, as service systems become increasingly…

统计方法学 · 统计学 2025-09-10 Shuze Chen , David Simchi-Levi , Chonghuan Wang

In online randomized experiments or A/B tests, accurate predictions of participant inclusion rates are of paramount importance. These predictions not only guide experimenters in optimizing the experiment's duration but also enhance the…

统计方法学 · 统计学 2024-02-06 Lorenzo Masoero , Mario Beraha , Thomas Richardson , Stefano Favaro

Over the past decade, most technology companies and a growing number of conventional firms have adopted online experimentation (or A/B testing) into their product development process. Initially, A/B testing was deployed as a static…

应用统计 · 统计学 2021-11-04 Jialiang Mao , Iavor Bojinov

Despite its importance to experimental design, statistical power (the probability that, given a real effect, an experiment will reject the null hypothesis) has largely been ignored by the NLP community. Underpowered experiments make it more…

计算与语言 · 计算机科学 2020-10-15 Dallas Card , Peter Henderson , Urvashi Khandelwal , Robin Jia , Kyle Mahowald , Dan Jurafsky

Online controlled experiment (also called A/B test or experiment) is the most important tool for decision-making at a wide range of data-driven companies like Microsoft, Google, Meta, etc. Metric computation is the core procedure for…

分布式、并行与集群计算 · 计算机科学 2024-09-27 Tao Xiong , Yong Wang

Online controlled experiments (A/B tests) are fundamental to data-driven decision-making in the digital economy. However, their real-world application is frequently compromised by two critical shortcomings: the use of statistically flawed…

应用统计 · 统计学 2025-09-30 Srijesh Pillai , Rajesh Kumar Chandrawat

Randomized experiments play a major role in data-driven decision making across many different fields and disciplines. In medicine, for example, randomized controlled trials (RCTs) are the backbone of clinical trial methodology for testing…

应用统计 · 统计学 2016-08-30 Andrew W. Correia

Machine learning models are being used extensively in many important areas, but there is no guarantee a model will always perform well or as its developers intended. Understanding the correctness of a model is crucial to prevent potential…

机器学习 · 计算机科学 2021-04-13 Huong Ha , Sunil Gupta , Santu Rana , Svetha Venkatesh

Adaptive experiments are used extensively in online platforms, healthcare and biotechnology, and a variety of other settings. In many of these applications, the main goal is not to precisely estimate a treatment effect, but to demonstrate…

A/B tests, also known as randomized controlled experiments (RCTs), are the gold standard for evaluating the impact of new policies, products, or decisions. However, these tests can be costly in terms of time and resources, potentially…

机器学习 · 统计学 2025-01-03 Shima Nassiri , Mohsen Bayati , Joe Cooprider

A/B testing refers to the statistical procedure of conducting an experiment to compare two treatments, A and B, applied to different testing subjects. It is widely used by technology companies such as Facebook, LinkedIn, and Netflix, to…

统计方法学 · 统计学 2026-05-12 Victoria Pokhiko , Qiong Zhang , Lulu Kang , D'arcy P. Mays

Randomized experiments (often known as "A/B tests") are widely used to evaluate product and service innovations. We study how to allocate limited experimentation resources across M concurrent experiments in an experiment-rich regime.…

统计方法学 · 统计学 2026-03-19 Fenghua Yang , Dae Woong Ham , Stefanus Jasin

A key trait of stochastic optimizers is that multiple runs of the same optimizer in attempting to solve the same problem can produce different results. As a result, their performance is evaluated over several repeats, or runs, on the…

机器学习 · 计算机科学 2026-05-18 Moslem Noori , Elisabetta Valiante , Thomas Van Vaerenbergh , Masoud Mohseni , Ignacio Rozada