中文
相关论文

相关论文: Setting the duration of online A/B experiments

200 篇论文

Online controlled experiments, such as A/B-tests, are commonly used by modern tech companies to enable continuous system improvements. Despite their paramount importance, A/B-tests are expensive: by their very definition, a percentage of…

机器学习 · 计算机科学 2024-01-09 Shubham Baweja , Neeti Pokharna , Aleksei Ustimenko , Olivier Jeunen

Many organizations utilize large-scale online controlled experiments (OCEs) to accelerate innovation. Having high statistical power to detect small differences between control and treatment accurately is critical, as even small changes in…

应用统计 · 统计学 2020-09-11 Ali Mahmoudzadeh , Sophia Liu , Sol Sadeghi , Paul Luo Li , Somit Gupta

In streaming platforms churn is extremely costly, yet A/B tests are typically evaluated using outcomes observed within a limited experimental horizon. Even when both short- and predicted long-term engagement metrics are considered, they may…

机器学习 · 计算机科学 2026-04-23 Dario Simionato , Andrea Tonon , Mingxue Wang , Weiguo Wang , Tong Gui , Xiaoyue Li

Constructing confidence intervals (CIs) for the average treatment effect (ATE) from patient records is crucial to assess the effectiveness and safety of drugs. However, patient records typically come from different hospitals, thus raising…

机器学习 · 计算机科学 2025-10-16 Yuxin Wang , Maresa Schröder , Dennis Frauen , Jonas Schweisthal , Konstantin Hess , Stefan Feuerriegel

Online controlled experiments, or A/B tests, are large-scale randomized trials in digital environments. This paper investigates the estimands of the difference-in-means estimator in these experiments, focusing on scenarios with repeated…

统计方法学 · 统计学 2024-11-12 Sebastian Ankargren , Mattias Frånberg , Mårten Schultzberg

Estimating the effects of long-term treatments through A/B testing is challenging. Treatments, such as updates to product functionalities, user interface designs, and recommendation algorithms, are intended to persist within the system for…

计量经济学 · 经济学 2025-12-30 Shan Huang , Chen Wang , Yuan Yuan , Jinglong Zhao , Brocco , Zhang

A/B tests serve the purpose of reliably identifying the effect of changes introduced in online services. It is common for online platforms to run a large number of simultaneous experiments by splitting incoming user traffic randomly in…

In adaptive clinical trials, the conventional confidence interval (CI) for a treatment effect is prone to undesirable properties such as undercoverage and potential inconsistency with the final hypothesis testing decision. Accordingly, as…

We study properties of confidence intervals (CIs) for the difference of two Bernoulli distributions' success parameters, $p_x - p_y$, in the case where the goal is to obtain a CI of a given half-width while minimizing sampling costs when…

统计方法学 · 统计学 2023-11-23 Ignacio Erazo , David Goldsman , Yajun Mei

While there exists a large amount of literature on the general challenges of and best practices for trustworthy online A/B testing, there are limited studies on sample size estimation, which plays a crucial role in trustworthy and efficient…

统计方法学 · 统计学 2023-08-21 Jing Zhou , Jiannan Lu , Anas Shallah

Sample size determination for cluster randomised trials (CRTs) is challenging as it requires robust estimation of the intra-cluster correlation coefficient (ICC). Typically, the sample size is chosen to provide a certain level of power to…

应用统计 · 统计学 2023-08-23 S. Faye Williamson , Svetlana V. Tishkovskaya , Kevin J. Wilson

This article develops new closed-form variance expressions for power analyses for commonly used difference-in-differences (DID) and comparative interrupted time series (CITS) panel data estimators. The main contribution is to incorporate…

统计方法学 · 统计学 2021-10-18 Peter Z. Schochet

The selection of the assumed effect size (AES) critically determines the duration of an experiment, and hence its accuracy and efficiency. Traditionally, experimenters determine AES based on domain knowledge. However, this method becomes…

机器学习 · 计算机科学 2025-04-15 Yu Liu , Runzhe Wan , James McQueen , Doug Hains , Jinxiang Gu , Rui Song

Recent years have seen the development of many novel scoring tools for disease prognosis and prediction. To become accepted for use in clinical applications, these tools have to be validated on external data. In practice, validation is…

统计方法学 · 统计学 2022-12-06 Matthias Schmid , Tim Friede , Nadja Klein , Leonie Weinhold

Companies offering web services routinely run randomized online experiments to estimate the causal impact associated with the adoption of new features and policies on key performance metrics of interest. These experiments are used to…

统计方法学 · 统计学 2023-07-13 Lorenzo Masoero , Doug Hains , James McQueen

Small sample sizes in clinical studies arises from factors such as reduced costs, limited subject availability, and the rarity of studied conditions. This creates challenges for accurately calculating confidence intervals (CIs) using the…

统计方法学 · 统计学 2025-11-11 Mulan Wu , Mengyu Xu , Dongyun Kim

Experimental testing is vital in the optimization of web applications, and as such A/B testing has been widely adopted as a methodology for determining optimal content for many web applications. While some testing platforms provide…

统计方法学 · 统计学 2017-10-04 Ian E. Fellows

Randomized experimentation (also known as A/B testing or bucket testing) is widely used in the internet industry to measure the metric impact obtained by different treatment variants. A/B tests identify the treatment variant showing the…

统计方法学 · 统计学 2020-12-23 Ye Tu , Kinjal Basu , Cyrus DiCiccio , Romil Bansal , Preetam Nandy , Padmini Jaikumar , Shaunak Chatterjee

We study nonasymptotic (finite-sample) confidence intervals for treatment effects in randomized experiments. In the existing literature, the effective sample sizes of nonasymptotic confidence intervals tend to be looser than the…

A generalization of the classical concordance correlation coefficient (CCC) is considered under a three-level design where multiple raters rate every subject over time, and each rater is rating every subject multiple times at each measuring…

统计方法学 · 统计学 2025-04-15 Soumya Sahu , Thomas Mathew , Dulal K. Bhaumik
‹ 上一页 1 2 3 10 下一页 ›