English
Related papers

Related papers: Setting the duration of online A/B experiments

200 papers

The determination of the sample size required by a crossover trial typically depends on the specification of one or more variance components. Uncertainty about the value of these parameters at the design stage means that there is often a…

Methodology · Statistics 2018-03-28 Michael Grayling , Adrian Mander , James Wason

Online controlled experiments (A/B tests) have become the gold standard for learning the impact of new product features in technology companies. Randomization enables the inference of causality from an A/B test. The randomized assignment…

Applications · Statistics 2022-12-20 Qike Li , Samir Jamkhande , Pavel Kochetkov , Pai Liu

Online controlled experiments, colloquially known as A/B-tests, are the bread and butter of real-world recommender system evaluation. Typically, end-users are randomly assigned some system variant, and a plethora of metrics are then…

Information Retrieval · Computer Science 2024-07-31 Olivier Jeunen , Shubham Baweja , Neeti Pokharna , Aleksei Ustimenko

Cluster randomized trials with measurements at baseline can improve power over post-test only designs by using difference in difference designs. However, subjects may be lost to follow-up between the baseline and follow-up periods. While…

Methodology · Statistics 2019-03-26 Jonathan Moyer , Ken Kleinman

Online A/B tests have become increasingly popular and important for social platforms. However, accurately estimating the global average treatment effect (GATE) has proven to be challenging due to network interference, which violates the…

Methodology · Statistics 2023-11-27 Qianyi Chen , Bo Li , Lu Deng , Yong Wang

A reasonable confidence interval should have a confidence coefficient no less than the given nominal level and a small expected length to reliably and accurately estimate the parameter of interest, and the bootstrap interval is considered…

Statistics Theory · Mathematics 2024-02-15 Weizhen Wang , Chongxiu Yu , Zhongzhan Zhang

Conformal prediction (CP) has become a cornerstone of distribution-free uncertainty quantification, conventionally evaluated by its coverage and interval length. This work critically examines the sufficiency of these standard metrics. We…

Machine Learning · Statistics 2026-01-30 Yizhou Min , Yizhou Lu , Lanqi Li , Zhen Zhang , Jiaye Teng

On-line experimentation (also known as A/B testing) has become an integral part of software development. To timely incorporate user feedback and continuously improve products, many software companies have adopted the culture of agile…

Applications · Statistics 2019-08-13 Yu Wang , Somit Gupta , Jiannan Lu , Ali Mahmoudzadeh , Sophia Liu

By employing various empirical estimators for the Mutual Information (MI) measure, we calculate and compare the estimates and their confidence intervals for both normal and non-normal bivariate data samples. We find that certain nonlinear…

Information Theory · Computer Science 2024-10-10 Theo Grigorenko , Leo Grigorenko

Objectives: Estimation of areas under receiver operating characteristic curves (AUCs) and their differences is a key task in diagnostic studies. We aimed to derive, evaluate, and implement simple sample size formulas for such studies with a…

Methodology · Statistics 2022-08-03 Di Shu , Guangyong Zou

Contextual sensing and delivery of digital interventions to improve health outcomes have gained significant traction in behavioral and psychiatric studies. Micro-randomized trials (MRTs) are a common experimental design for obtaining…

Methodology · Statistics 2025-04-01 Jieru Shi , Zhenke Wu , Walter Dempsey

A/B tests are the gold standard for evaluating digital experiences on the web. However, traditional "fixed-horizon" statistical methods are often incompatible with the needs of modern industry practitioners as they do not permit continuous…

Composite binary endpoints are increasingly used as primary endpoints in clinical trials. When designing a trial, it is crucial to determine the appropriate sample size for testing the statistical differences between treatment groups for…

Applications · Statistics 2019-01-15 Marta Bofill Roig , Guadalupe Gómez Melis

This paper investigates estimating the variance of a temporal-difference learning agent's update target. Most reinforcement learning methods use an estimate of the value function, which captures how good it is for the agent to be in a…

Artificial Intelligence · Computer Science 2018-02-15 Craig Sherstan , Brendan Bennett , Kenny Young , Dylan R. Ashley , Adam White , Martha White , Richard S. Sutton

Stratification is commonly employed in clinical trials to reduce the chance covariate imbalances and increase the precision of the treatment effect estimate. We propose a general framework for constructing the confidence interval (CI) for a…

Methodology · Statistics 2021-10-26 Yongqiang Tang

Multi-regional clinical trials (MRCTs) have become common practice for drug development and global registration. Once overall significance is established, demonstrating regional consistency is critical for local health authorities. Methods…

Methodology · Statistics 2025-08-14 Xinru Ren , Jin Xu

When designing experimental studies with human participants, experimenters must decide how many trials each participant will complete, as well as how many participants to test. Most discussion of statistical power (the ability of a study…

Neurons and Cognition · Quantitative Biology 2021-08-30 Daniel H. Baker , Greta Vilidaite , Freya A. Lygo , Anika K. Smith , Tessa R. Flack , Andre D. Gouws , Timothy J. Andrews

In this paper, we consider an experimental setting where units enter the experiment sequentially. Our goal is to form stopping rules which lead to estimators of treatment effects with a given precision. We propose a fixed-width confidence…

Statistics Theory · Mathematics 2024-05-07 Mattias Nordin , Mårten Schultzberg

We consider the design of a two-arm superiority cluster randomised controlled trial (RCT) with a continuous outcome. We detail Bayesian inference for the analysis of the trial using a linear mixed-effects model. The treatment is compared to…

Methodology · Statistics 2022-08-29 Kevin J Wilson

Receiver operating characteristic (ROC) curves are widely used as a measure of accuracy of diagnostic tests and can be summarized using the area under the ROC curve (AUC). Often, it is useful to construct a confidence intervals for the AUC,…

Applications · Statistics 2018-04-18 Hunyong Cho , Gregory J. Matthews , Ofer Harel
‹ Prev 1 3 4 5 6 7 10 Next ›