中文
相关论文

相关论文: A Common Misassumption in Online Experiments with …

200 篇论文

Current approaches to A/B testing in networks focus on limiting interference, the concern that treatment effects can "spill over" from treatment nodes to control nodes and lead to biased causal effect estimation. Prominent methods for…

机器学习 · 计算机科学 2020-04-16 Zahra Fatemi , Elena Zheleva

It has been recently shown in the literature that the sample averages from online learning experiments are biased when used to estimate the mean reward. To correct the bias, off-policy evaluation methods, including importance sampling and…

机器学习 · 计算机科学 2021-12-02 Ningyuan Chen , Xuefeng Gao , Yi Xiong

Online controlled experiments, also known as A/B testing, are the digital equivalent of randomized controlled trials for estimating the impact of marketing campaigns on website visitors. Stratified sampling is a traditional technique for…

Methods that infer causal dependence from observational data are central to many areas of science, including medicine, economics, and the social sciences. A variety of theoretical properties of these methods have been proven, but empirical…

统计方法学 · 统计学 2021-07-08 Amanda Gentzel , Purva Pruthi , David Jensen

Digital technology organizations routinely use online experiments (e.g. A/B tests) to guide their product and business decisions. In e-commerce, we often measure changes to transaction- or item-based business metrics such as Average Basket…

应用统计 · 统计学 2023-04-18 C. H. Bryan Liu , Emma J. McCoy

A/B testing has become a gold standard for modern technological companies to conduct policy evaluation. Yet, its application to time series experiments, where policies are sequentially assigned over time, remains challenging. Existing…

机器学习 · 计算机科学 2026-02-03 Xiangkun Wu , Qianglin Wen , Yingying Zhang , Hongtu Zhu , Ting Li , Chengchun Shi

A standard assumption in machine learning is the exchangeability of data, which is equivalent to assuming that the examples are generated from the same probability distribution independently. This paper is devoted to testing the assumption…

机器学习 · 计算机科学 2012-06-29 Valentina Fedorova , Alex Gammerman , Ilia Nouretdinov , Vladimir Vovk

Online A/B tests have become increasingly popular and important for social platforms. However, accurately estimating the global average treatment effect (GATE) has proven to be challenging due to network interference, which violates the…

统计方法学 · 统计学 2023-11-27 Qianyi Chen , Bo Li , Lu Deng , Yong Wang

Randomized Controlled Trials (RCT) are the current gold standards to empirically measure the effect of a new drug. However, they may be of limited size and resorting to complementary non-randomized data, referred to as observational, is…

统计方法学 · 统计学 2025-06-11 Ahmed Boughdiri , Julie Josse , Erwan Scornet

Online controlled experiments, now commonly known as A/B testing, are crucial to causal inference and data driven decision making in many internet based businesses. While a simple comparison between a treatment (the feature under test) and…

应用统计 · 统计学 2015-01-05 Yu Guo , Alex Deng

Hybrid randomized controlled trials (hybrid RCTs) integrate external control data, such as historical or concurrent data, with data from randomized trials. While numerous frequentist and Bayesian methods, such as the test-then-pool and…

统计方法学 · 统计学 2025-10-07 Han Chang Chiam , Franz König , Martin Posch

With software systems becoming increasingly pervasive and autonomous, our ability to test for their quality is severely challenged. Many systems are called to operate in uncertain and highly-changing environment, not rarely required to make…

软件工程 · 计算机科学 2024-03-21 Luca Giamattei , Roberto Pietrantuono , Stefano Russo

In this work, we proposed a novel inferential procedure assisted by machine learning based adjustment for randomized control trials. The method was developed under the Rosenbaum's framework of exact tests in randomized experiments with…

统计方法学 · 统计学 2024-07-23 Han Yu , Alan D. Hutson , Xiaoyi Ma

A central obstacle in the objective assessment of treatment effect (TE) estimators in randomized control trials (RCTs) is the lack of ground truth (or validation set) to test their performance. In this paper, we propose a novel…

In recent years, attention has increasingly focused on enhancing user satisfaction with user interfaces, spanning both mobile applications and websites. One fundamental aspect of human-machine interaction is the concept of web usability. In…

Testing practices within the machine learning (ML) community have centered around assessing a learned model's predictive performance measured against a test dataset, often drawn from the same distribution as the training dataset. While…

机器学习 · 计算机科学 2021-12-07 Negar Rostamzadeh , Ben Hutchinson , Christina Greer , Vinodkumar Prabhakaran

Safety-critical robot systems need thorough testing to expose design flaws and software bugs which could endanger humans. Testing in simulation is becoming increasingly popular, as it can be applied early in the development process and does…

机器人学 · 计算机科学 2023-11-07 Tom P. Huck , Martin Kaiser , Constantin Cronrath , Bengt Lennartson , Torsten Kröger , Tamim Asfour

Tech companies (e.g., Google or Facebook) often use randomized online experiments and/or A/B testing primarily based on the average treatment effects to compare their new product with an old one. However, it is also critically important to…

统计方法学 · 统计学 2021-11-09 Chengchun Shi , Shikai Luo , Hongtu Zhu , Rui Song

Typically, a randomized experiment is designed to test a hypothesis about the average treatment effect and sometimes hypotheses about treatment effect variation. The results of such a study may then be used to inform policy and practice for…

统计方法学 · 统计学 2026-05-01 Elizabeth Tipton , Michalis Mamakos

Risk assessment of a robot in controlled environments, such as laboratories and proving grounds, is a common means to assess, certify, validate, verify, and characterize the robots' safety performance before, during, and even after their…

机器人学 · 计算机科学 2025-01-29 Linda Capito , Guillermo A. Castillo , Bowen Weng