中文
相关论文

相关论文: Test & Roll: Profit-Maximizing A/B Tests

200 篇论文

Online media platforms often need to measure how frequently users are exposed to specific content attributes in order to evaluate trade-offs in A/B experiments. A direct approach is to sample content, label it using a high-quality rubric…

应用统计 · 统计学 2026-02-19 Zehao Xu , Tony Paek , Kevin O'Sullivan , Attila Dobi

Empirical economic studies often involve multiple propositions or hypotheses, with researchers aiming to assess both the collective and individual evidence against these propositions or hypotheses. To rigorously assess this evidence,…

计量经济学 · 经济学 2024-08-26 Zeng-Hua Lu

Efficient methods to evaluate new algorithms are critical for improving interactive bandit and reinforcement learning systems such as recommendation systems. A/B tests are reliable, but are time- and money-consuming, and entail a risk of…

机器学习 · 计算机科学 2021-08-04 Yusuke Narita , Shota Yasui , Kohei Yata

This paper studies the identification, estimation, and hypothesis testing problem in complete and incomplete economic models with testable assumptions. Testable assumptions ($A$) give strong and interpretable empirical content to the models…

计量经济学 · 经济学 2022-03-11 Moyu Liao

In this paper we have updated the hypothesis testing framework by drawing upon modern computational power and classification models from machine learning. We show that a simple classification algorithm such as a boosted decision stump can…

计量经济学 · 经济学 2021-03-03 Gary Cornwall , Jeff Chen , Beau Sauley

A common dilemma encountered by many upon implementing an optimization method or experiment, whether it be a reinforcement learning algorithm, or A/B testing, is deciding on what metric to optimize for. Very often short-term metrics, which…

应用统计 · 统计学 2019-06-17 Yoni Schamroth , Liron Gat Kahlon , Boris Rabinovich , David Steinberg

Multiple hypothesis testing practices vary widely, without consensus on which are appropriate when. This paper provides an economic foundation for these practices designed to capture leading examples, such as regulatory approval on the…

综合经济学 · 经济学 2026-02-19 Davide Viviano , Kaspar Wuthrich , Paul Niehaus

There are multiple testing methods to ascertain an infection in an individual and they vary in their performances, cost and delay. Unfortunately, better performing tests are sometimes costlier and time consuming and can only be done for a…

社会与信息网络 · 计算机科学 2021-06-17 Harish Sasikumar , Manoj Varma

Empirical researchers often trim observations with small denominator A when they estimate moments of the form E[B/A]. Large trimming is a common practice to mitigate variance, but it incurs large trimming bias. This paper provides a novel…

统计方法学 · 统计学 2021-01-12 Yuya Sasaki , Takuya Ura

Experimental design has emerged as a powerful approach for improving the sample efficiency of A/B testing, yet existing designs rely critically on correctly specified models. We study robust sequential experimental design under model…

机器学习 · 统计学 2026-05-14 Qianglin Wen , Xiangkun Wu , Chengchun Shi , Ting Li , Niansheng Tang , Yingying Zhang , Hongtu Zhu

Underpowered studies (below 50% power) suffer from the winner's curse: A statistically significant positive estimate must exaggerate the true treatment effect to meet the significance threshold. A study by Dipayan Biswas, Annika Abell, and…

Preferably in two- or three-arm randomized clinical trials, a few (2,3) correlated multiple primary endpoints are considered. In addition to the closed testing principle based on different global tests, two max(maxT) tests are compared with…

应用统计 · 统计学 2021-03-16 Ludwig A. Hothorn , Siegfried Kropf

In this paper, we address the fundamental statistical question: how can you assess the power of an A/B test when the units in the study are exposed to interference? This question is germane to many scientific and industrial practitioners…

社会与信息网络 · 计算机科学 2017-10-12 James D. Wilson , David T. Uminsky

Accurate estimation of treatment effects in online A/B testing is challenging with zero-inflated and skewed metrics. Traditional tests, like Welch's t-test, often lack sensitivity with heavy-tailed data due to their reliance on means, as…

统计方法学 · 统计学 2025-10-07 Kevin Charette , Tristan Boudreault

Online experiments such as Randomised Controlled Trials (RCTs) or A/B-tests are the bread and butter of modern platforms on the web. They are conducted continuously to allow platforms to estimate the causal effect of replacing system…

机器学习 · 计算机科学 2023-04-24 Olivier Jeunen

Comparison Lift is an experimentation-as-a-service (EaaS) application for testing online advertising audiences and creatives at JD.com. Unlike many other EaaS tools that focus primarily on fixed sample A/B testing, Comparison Lift deploys a…

机器学习 · 计算机科学 2020-09-18 Tong Geng , Xiliang Lin , Harikesh S. Nair , Jun Hao , Bin Xiang , Shurui Fan

Trial-offer markets, where customers can sample a product before deciding whether to buy it, are ubiquitous in the online experience. Their static and dynamic properties are often studied by assuming that consumers follow a multinomial…

计算机科学与博弈论 · 计算机科学 2016-10-07 Pascal Van Hentenryck , Alvaro Flores , Gerardo Berbeglia

In recommender systems, online A/B testing is a crucial method for evaluating the performance of different models. However, conducting online A/B testing often presents significant challenges, including substantial economic costs, user…

While developments in machine learning led to impressive performance gains on big data, many human subjects data are, in actuality, small and sparsely labeled. Existing methods applied to such data often do not easily generalize to…

机器学习 · 计算机科学 2023-04-04 Julie Jiang , Kristina Lerman , Emilio Ferrara

Direct marketers use target models in order to minimize the spreading loss of sales efforts. The application of target models has become more widespread with the increasing range of sales efforts. Target models are relevant for offline…

应用统计 · 统计学 2010-07-08 Joerg Dubiel