中文
相关论文

相关论文: Size matters? Or not: A/B testing with limited sam…

200 篇论文

Large language models (LLMs) have shown remarkable adaptability to diverse tasks, by leveraging context prompts containing instructions, or minimal input-output examples. However, recent work revealed they also exhibit label bias -- an…

计算与语言 · 计算机科学 2024-05-07 Yuval Reif , Roy Schwartz

A/B tests have been widely adopted across industries as the golden rule that guides decision making. However, the long-term true north metrics we ultimately want to drive through A/B test may take a long time to mature. In these situations,…

应用统计 · 统计学 2021-06-04 Weitao Duan , Shan Ba , Chunzhe Zhang

Adaptive experimental design (AED) methods are increasingly being used in industry as a tool to boost testing throughput or reduce experimentation cost relative to traditional A/B/N testing methods. However, the behavior and guarantees of…

机器学习 · 计算机科学 2024-09-19 Tanner Fiez , Houssam Nassif , Yu-Cheng Chen , Sergio Gamez , Lalit Jain

Calibration weighting has been widely used to correct selection biases in non-probability sampling, missing data, and causal inference. The main idea is to calibrate the biased sample to the benchmark by adjusting the subject weights.…

统计方法学 · 统计学 2023-05-30 Chenyin Gao , Shu Yang , Jae Kwang Kim

Positivity violations, which occur when some subgroups either always or never receive a treatment of interest, pose significant challenges for causal effect estimation with observational data. Recent balancing weight methods have proved to…

统计方法学 · 统计学 2025-12-17 Martha Barnard , Jared D. Huling , Julian Wolfson

Innovations across science and industry are evaluated using randomized trials (a.k.a. A/B tests). While simple and robust, such static designs are inefficient or infeasible for testing many hypotheses. Adaptive designs can greatly improve…

机器学习 · 计算机科学 2024-08-09 Jimmy Wang , Ethan Che , Daniel R. Jiang , Hongseok Namkoong

Tech companies (e.g., Google or Facebook) often use randomized online experiments and/or A/B testing primarily based on the average treatment effects to compare their new product with an old one. However, it is also critically important to…

统计方法学 · 统计学 2021-11-09 Chengchun Shi , Shikai Luo , Hongtu Zhu , Rui Song

Experimental testing is vital in the optimization of web applications, and as such A/B testing has been widely adopted as a methodology for determining optimal content for many web applications. While some testing platforms provide…

统计方法学 · 统计学 2017-10-04 Ian E. Fellows

Covariate balancing is a popular technique for controlling confounding in observational studies. It finds weights for the treatment group which are close to uniform, but make the group's covariate means (approximately) equal to those of the…

统计方法学 · 统计学 2025-03-07 Shiva Kaul , Min-Gyu Kim

In observational studies of treatment effects, matched samples are created so treated and control groups are similar in terms of observable covariates. Traditionally such matched samples consist of matched pairs. If a pair match fails to…

统计方法学 · 统计学 2014-10-22 Luke Keele , Sam Pimentel , Frank Yoon

Digital firms routinely run many online experiments on shared user populations. When product decisions are compositional, such as combinations of interface elements, flows, messages, or incentives, the number of feasible interventions grows…

机器学习 · 统计学 2026-04-13 Xin Wen , Xi Chen , Will Wei Sun , Yichen Zhang

The automotive domain is shifting to software-centric development to meet regulation, market pressure, and feature velocity. This shift increases embedded systems' complexity and strains testing capacity. Despite relevant standards, a…

软件工程 · 计算机科学 2026-01-09 Denesa Zyberaj , Pascal Hirmer , Marco Aiello , Stefan Wagner

Online experiments such as Randomised Controlled Trials (RCTs) or A/B-tests are the bread and butter of modern platforms on the web. They are conducted continuously to allow platforms to estimate the causal effect of replacing system…

机器学习 · 计算机科学 2023-04-24 Olivier Jeunen

Online experiments (A/B tests) are widely regarded as the gold standard for evaluating recommender system variants and guiding launch decisions. However, a variety of biases can distort the results of the experiment and mislead…

信息检索 · 计算机科学 2025-09-03 Chen Zheng , Zhenyu Zhao

A/B testing is widexly used in the industry to optimize customer facing websites. Many companies employ experimentation specialists to facilitate and improve the process of A/B testing. Here, we present the application of A/B testing to…

信息检索 · 计算机科学 2024-06-25 Melanie J. I. Müller

User-randomized A/B testing has emerged as the gold standard for online experimentation. However, when this kind of approach is not feasible due to legal, ethical or practical considerations, experimenters have to consider alternatives like…

统计方法学 · 统计学 2025-06-17 Paul Missault , Lorenzo Masoero , Christian Delbé , Thomas Richardson , Guido Imbens

Agile methodologies have gained significant traction in the software development industry, promising increased flexibility and responsiveness to changing requirements. However, their applicability to safety-critical systems, particularly in…

软件工程 · 计算机科学 2024-09-20 Mehrnoosh Askarpour , Sahar Kokaly , Ramesh S

A challenge that machine learning practitioners in the industry face is the task of selecting the best model to deploy in production. As a model is often an intermediate component of a production system, online controlled experiments such…

Compute and memory constraints have historically prevented traffic simulation software users from fully utilizing the predictive models underlying them. When calibrating car-following models, particularly, accommodations have included 1)…

机器学习 · 统计学 2019-08-08 Franklin Abodo , Andrew Berthaume , Stephen Zitzow-Childs , Leonardo Bobadilla

Software companies have widely used online A/B testing to evaluate the impact of a new technology by offering it to groups of users and comparing it against the unmodified product. However, running online A/B testing needs not only efforts…

软件工程 · 计算机科学 2024-08-12 Jie JW Wu