English
Related papers

Related papers: It's About Time: What A/B Test Metrics Estimate

200 papers

Online experiments are the gold standard for evaluating impact on user experience and accelerating innovation in software. However, since experiments are typically limited in duration, observed treatment effects are not always permanently…

Human-Computer Interaction · Computer Science 2021-02-26 Soheil Sadeghi , Somit Gupta , Stefan Gramatovici , Jiannan Lu , Hao Ai , Ruhan Zhang

Effective decision making from randomised controlled clinical trials relies on robust interpretation of the numerical results. However, the language we use to describe clinical trials can cause confusion both in trial design and in…

A/B tests, also known as randomized controlled experiments (RCTs), are the gold standard for evaluating the impact of new policies, products, or decisions. However, these tests can be costly in terms of time and resources, potentially…

Machine Learning · Statistics 2025-01-03 Shima Nassiri , Mohsen Bayati , Joe Cooprider

Achieving error rates that meet or exceed the fault-tolerance threshold is a central goal for quantum computing experiments, and measuring these error rates using randomized benchmarking is now routine. However, direct comparison between…

Quantum Physics · Physics 2016-10-26 Richard Kueng , David M. Long , Andrew C. Doherty , Steven T. Flammia

Before the availability of large scale fault-tolerant quantum devices, one has to find ways to make the most of current noisy intermediate-scale quantum devices. One possibility is to seek smaller repetitive hybrid quantum-classical tasks…

Quantum Physics · Physics 2023-04-12 Teiko Heinosaari , Daniel Reitzner , Alessandro Toigo

We propose a classification of measurement apparatuses based on their reliability and accessibility. Our notion of reliability parameterises the possibility of getting unexpected wrong results when using the apparatus in a given time…

Quantum Physics · Physics 2024-07-26 Nicola Pranzini , Paola Verrucchi

A challenge that machine learning practitioners in the industry face is the task of selecting the best model to deploy in production. As a model is often an intermediate component of a production system, online controlled experiments such…

Machine Learning · Statistics 2021-05-31 Zhenwen Dai , Praveen Chandar , Ghazal Fazelnia , Ben Carterette , Mounia Lalmas-Roelleke

In this paper, we investigate the problem of assessing statistical methods and effectively summarizing results from simulations. Specifically, we consider problems of the type where multiple methods are compared on a reasonably large test…

Applications · Statistics 2015-10-07 Abigail Arnold , Jason Loeppky

eBay's experimentation platform runs hundreds of A/B tests on any given day. The platform integrates with the tracking infrastructure and customer experience servers, provides the sampling service for experiments, and has the responsibility…

Applications · Statistics 2023-03-10 Keyu Nie , Zezhong Zhang , Bingquan Xu , Tao Yuan

How should researchers analyze randomized experiments in which the main outcome is latent and measured in multiple ways but each measure contains some degree of error? We first identify a critical study-specific noncomparability problem in…

Econometrics · Economics 2026-01-13 Jiawei Fu , Donald P. Green

A method is proposed that allows one to infer the sum of the values of an observable taken during contacts with a pointer state. Hereby the state of the pointer is updated while contacted with the system and remains unchanged between…

Statistical Mechanics · Physics 2020-11-06 Juzar Thingna , Peter Talkner

In the era of large-scale AI deployment and high-stakes clinical trials, adaptive experimentation faces a ``trilemma'' of conflicting objectives: minimizing cumulative regret (welfare loss during the experiment), maximizing the estimation…

Methodology · Statistics 2026-02-13 Jiachun Li , Kaining Shi , David Simchi-Levi

Digital firms routinely run many online experiments on shared user populations. When product decisions are compositional, such as combinations of interface elements, flows, messages, or incentives, the number of feasible interventions grows…

Machine Learning · Statistics 2026-04-13 Xin Wen , Xi Chen , Will Wei Sun , Yichen Zhang

Though it has been recognized that recommending serendipitous (i.e., surprising and relevant) items can be helpful for increasing users' satisfaction and behavioral intention, how to measure serendipity in the offline environment is still…

Human-Computer Interaction · Computer Science 2020-04-23 Li Chen , Ningxia Wang , Yonghua Yang , Keping Yang , Quan Yuan

Mutual information is a widely-used information theoretic measure to quantify the amount of association between variables. It is used extensively in many applications such as image registration, diagnosis of failures in electrical machines,…

Computation · Statistics 2021-08-21 Luai Al-Labadi , Forough Fazeli-Asl , Zahra Saberi

There are a number of measures of direct and indirect effects in the literature. They are suitable in some cases and unsuitable in others. We describe a case where the existing measures are unsuitable and propose new suitable ones. We also…

Methodology · Statistics 2023-10-12 Jose M. Peña

In this paper, we examine the biases that arise when firms run A/B tests on continuous parameters to estimate global treatment effects on performance metrics of interest; we particularly focus on price experiments to measure the price…

Methodology · Statistics 2026-01-22 Ramesh Johari , Orrie B. Page , Gabriel Y. Weintraub

Online advertisements have become one of today's most widely used tools for enhancing businesses partly because of their compatibility with A/B testing. A/B testing allows sellers to find effective advertisement strategies such as ad…

Machine Learning · Computer Science 2020-10-22 Akira Matsui , Daisuke Moriwaki

According to the dominant view, time in perceptual decision making is used for integrating new sensory evidence. Based on a probabilistic framework, we investigated the alternative hypothesis that time is used for gradually refining an…

Neurons and Cognition · Quantitative Biology 2015-02-12 Máté Lengyel , Ádám Koblinger , Marjena Popović , József Fiser

We present a sample path dependent measure of causal influence between two time series. The proposed measure is a random variable whose expected sum is the directed information. A realization of the proposed measure may be used to identify…

Information Theory · Computer Science 2018-10-15 Gabriel Schamberg , Todd P. Coleman