English
Related papers

Related papers: It's About Time: What A/B Test Metrics Estimate

200 papers

Selecting the optimal recommender via online exploration-exploitation is catching increasing attention where the traditional A/B testing can be slow and costly, and offline evaluations are prone to the bias of history data. Finding the…

Information Retrieval · Computer Science 2022-03-29 Da Xu , Chuanwei Ruan , Evren Korpeoglu , Sushant Kumar , Kannan Achan

Predicting the timing and occurrence of events is a major focus of data science applications, especially in the context of biomedical research. Performance for models estimating these outcomes, often referred to as time-to-event or survival…

Methodology · Statistics 2024-06-07 Ying Jin , Andrew Leroux

A/B testing is ubiquitous within the machine learning and data science operations of internet companies. Generically, the idea is to perform a statistical test of the hypothesis that a new feature is better than the existing platform---for…

Statistics Theory · Mathematics 2017-10-11 David Goldberg , James E. Johndrow

We consider the problem of combining data from observational and experimental sources to make causal conclusions. This problem is increasingly relevant, as the modern era has yielded passive collection of massive observational datasets in…

Methodology · Statistics 2020-05-19 Evan Rosenman , Guillaume Basse , Art Owen , Michael Baiocchi

With the growth of interpreting technologies, from remote interpreting and Computer-Aided Interpreting to automated speech translation and interpreting avatars, there is now a high demand for ways to quickly and efficiently measure the…

Computation and Language · Computer Science 2026-01-12 Jonathan Downie , Joss Moorkens

Automatic evaluation metrics capable of replacing human judgments are critical to allowing fast development of new methods. Thus, numerous research efforts have focused on crafting such metrics. In this work, we take a step back and analyze…

Computation and Language · Computer Science 2022-10-10 Pierre Colombo , Maxime Peyrard , Nathan Noiry , Robert West , Pablo Piantanida

This paper studies the measurement of advertising effects on online platforms when parallel experimentation occurs, that is, when multiple advertisers experiment concurrently. It provides a framework that makes precise how parallel…

General Economics · Economics 2024-03-01 Caio Waisman , Navdeep S. Sahni , Harikesh S. Nair , Xiliang Lin

The quality of a summarization evaluation metric is quantified by calculating the correlation between its scores and human annotations across a large number of summaries. Currently, it is unclear how precise these correlation estimates are,…

Computation and Language · Computer Science 2021-07-28 Daniel Deutsch , Rotem Dror , Dan Roth

A primary difficulty with unsupervised discovery of structure in large data sets is a lack of quantitative evaluation criteria. In this work, we propose and investigate several metrics for evaluating and comparing generative models of…

Machine Learning · Computer Science 2020-07-27 Daniel Jiwoong Im , Iljung Kwak , Kristin Branson

Experimentation platforms in industry must often deal with customer trust issues. Platforms must prove the validity of their claims as well as catch issues that arise. As a central quantity estimated by experimentation platforms, the…

Methodology · Statistics 2025-11-21 Kedar Karhadkar , Jack Klys , Daniel Ting , Artem Vorozhtsov , Houssam Nassif

Automated metrics for Machine Translation have made significant progress, with the goal of replacing expensive and time-consuming human evaluations. These metrics are typically assessed by their correlation with human judgments, which…

Computation and Language · Computer Science 2024-12-31 Pius von Däniken , Jan Deriu , Mark Cieliebak

Human-robot collaboration enables highly adaptive co-working. The variety of resulting workflows makes it difficult to measure metrics as, e.g. makespans or idle times for multiple systems and tasks in a comparable manner. This issue can be…

Robotics · Computer Science 2024-11-15 Jonathan Hümmer , Dominik Riedelbauch , Dominik Henrich

Monitoring is an important part of the verification toolbox, in particular in situations where exhaustive verification using, e.g., model-checking is infeasible. The goal of online monitoring is to determine the satisfaction or violation of…

Formal Languages and Automata Theory · Computer Science 2025-10-02 Thomas M. Grosen , Sean Kauffman , Kim G. Larsen , Martin Zimmermann

We introduce the quantum indirect estimation theory, which provides a general framework to address the problem of which ensemble averages can be estimated by means of an available set of measuring apparatuses, e. g. estimate the ensemble…

Quantum Physics · Physics 2008-06-17 Giacomo Mauro D'Ariano , Paolo Perinotti , Massimiliano Federico Sacchi

The determination of the sample size required by a crossover trial typically depends on the specification of one or more variance components. Uncertainty about the value of these parameters at the design stage means that there is often a…

Methodology · Statistics 2018-03-28 Michael Grayling , Adrian Mander , James Wason

The standard A/B testing approaches are mostly based on t-test in large scale industry applications. These standard approaches however suffers from low statistical power in business settings, due to nature of small sample-size or…

Methodology · Statistics 2025-12-30 Changshuai Wei , Phuc Nguyen , Benjamin Zelditch , Joyce Chen

Randomized experimentation (also known as A/B testing or bucket testing) is widely used in the internet industry to measure the metric impact obtained by different treatment variants. A/B tests identify the treatment variant showing the…

A/B testing is an important decision making tool in product development because can provide an accurate estimate of the average treatment effect of a new features, which allows developers to understand how the business impact of new changes…

Applications · Statistics 2019-03-22 Guillaume Saint-Jacques , James Eric Sorenson , Nanyu Chen , Ya Xu

Several problems in statistics involve the combination of high-variance unbiased estimators with low-variance estimators that are only unbiased under strong assumptions. A notable example is the estimation of causal effects while combining…

Methodology · Statistics 2023-05-25 Michael Oberst , Alexander D'Amour , Minmin Chen , Yuyan Wang , David Sontag , Steve Yadlowsky

Over the past decade, most technology companies and a growing number of conventional firms have adopted online experimentation (or A/B testing) into their product development process. Initially, A/B testing was deployed as a static…

Applications · Statistics 2021-11-04 Jialiang Mao , Iavor Bojinov