中文
相关论文

相关论文: Powerful A/B-Testing Metrics and Where to Find The…

200 篇论文

Offline evaluations of recommender systems attempt to estimate users' satisfaction with recommendations using static data from prior user interactions. These evaluations provide researchers and developers with first approximations of the…

信息检索 · 计算机科学 2020-01-28 Mucun Tian , Michael D. Ekstrand

There are many offline metrics that can be used as a reference for evaluation and optimization of the performance of recommender systems. Hybrid recommendation approaches are commonly used to improve some of those metrics by combining…

信息检索 · 计算机科学 2019-01-09 Andres Ferraro , Dmitry Bogdanov , Kyumin Choi , Xavier Serra

We consider the estimation of heterogeneous treatment effects with arbitrary machine learning methods in the presence of unobserved confounders with the aid of a valid instrument. Such settings arise in A/B tests with an intent-to-treat…

计量经济学 · 经济学 2019-06-07 Vasilis Syrgkanis , Victor Lei , Miruna Oprescu , Maggie Hei , Keith Battocchi , Greg Lewis

Heavy-tailed metrics are common and often critical to product evaluation in the online world. While we may have samples large enough for Central Limit Theorem to kick in, experimentation is challenging due to the wide confidence interval of…

应用统计 · 统计学 2019-05-23 Jason , Wang , Pauline Burke

A/B testing is critical for modern technological companies to evaluate the effectiveness of newly developed products against standard baselines. This paper studies optimal designs that aim to maximize the amount of information obtained from…

统计方法学 · 统计学 2023-11-07 Ting Li , Chengchun Shi , Jianing Wang , Fan Zhou , Hongtu Zhu

The evaluation of recommender system fairness has become increasingly important, especially with recent legislation that emphasises the development of fair and responsible artificial intelligence. This has led to the emergence of various…

信息检索 · 计算机科学 2026-04-29 Theresia Veronika Rampisela

Alternative metrics (aka altmetrics) are gaining increasing interest in the scientometrics community as they can capture both the volume and quality of attention that a research work receives online. Nevertheless, there is limited knowledge…

The rollout of new versions of a feature in modern applications is a manual multi-stage process, as the feature is released to ever larger groups of users, while its performance is carefully monitored. This kind of A/B testing is…

机器学习 · 计算机科学 2018-05-29 Andrés Muñoz Medina , Sergei Vassilvitskii , Dong Yin

Design of experiments and estimation of treatment effects in large-scale networks, in the presence of strong interference, is a challenging and important problem. Most existing methods' performance deteriorates as the density of the network…

统计方法学 · 统计学 2020-12-15 Preetam Nandy , Kinjal Basu , Shaunak Chatterjee , Ye Tu

Gathering observational data for medical decision-making often involves uncertainties arising from both type I (false positive)and type II (false negative) errors. In this work, we develop a statistical model to study how medical…

应用统计 · 统计学 2025-10-21 Lucas Böttcher , Maria R. D'Orsogna , Tom Chou

Online controlled experimentation is widely adopted for evaluating new features in the rapid development cycle for web products and mobile applications. Measurement of the overall experiment sample is a common practice to quantify the…

人机交互 · 计算机科学 2022-01-27 Zhenyu Zhao , Yan He , Miao Chen

Evaluating retrieval-ranking systems is crucial for developing high-performing models. While online A/B testing is the gold standard, its high cost and risks to user experience require effective offline methods. However, relying on…

信息检索 · 计算机科学 2025-04-08 Seyedeh Baharan Khatami , Sayan Chakraborty , Ruomeng Xu , Babak Salimi

We study the problem of high-dimensional robust mean estimation in an online setting. Specifically, we consider a scenario where $n$ sensors are measuring some common, ongoing phenomenon. At each time step $t=1,2,\ldots,T$, the $i^{th}$…

机器学习 · 计算机科学 2023-10-26 Daniel M. Kane , Ilias Diakonikolas , Hanshen Xiao , Sihan Liu

A/B testing plays a central role in data-driven product development, guiding launch decisions for new features and designs. However, treatment effect estimates are often noisy due to short horizons, early stopping, and slowly accumulating…

统计方法学 · 统计学 2025-11-27 Xinran Li

Click-through rate (CTR) prediction is a crucial task in online advertising to recommend products that users are likely to be interested in. To identify the best-performing models, rigorous model evaluation is necessary. Offline…

信息检索 · 计算机科学 2024-06-27 Ramazan Tarik Turksoy , Beyza Turkmen

Ordinal user-provided ratings across multiple items are frequently encountered in both scientific and commercial applications. Whilst recommender systems are known to do well on these type of data from a predictive point of view, their…

统计方法学 · 统计学 2025-03-05 Sjoerd Hermes

Altmetrics are tools for measuring the impact of research beyond scientific communities. In general, they measure online mentions of scholarly outputs, such as on online social networks, blogs, and news sites. Some stakeholders in higher…

数字图书馆 · 计算机科学 2021-12-20 Grischa Fraumann

In streaming platforms churn is extremely costly, yet A/B tests are typically evaluated using outcomes observed within a limited experimental horizon. Even when both short- and predicted long-term engagement metrics are considered, they may…

机器学习 · 计算机科学 2026-04-23 Dario Simionato , Andrea Tonon , Mingxue Wang , Weiguo Wang , Tong Gui , Xiaoyue Li

Online media offers opportunities to marketers to deliver brand messages to a large audience. Advertising technology platforms enables the advertisers to find the proper group of audiences and deliver ad impressions to them in real time.…

人工智能 · 计算机科学 2016-02-24 Bowen Zhou , Shahriar Shariat

Online A/B experiments generate millions of user-activity records each day, yet experimenters need timely forecasts to guide roll-outs and safeguard user experience. Motivated by the problem of activity prediction for A/B tests at Amazon,…

应用统计 · 统计学 2025-05-27 Mario Beraha , Lorenzo Masoero , Stefano Favaro , Thomas S. Richardson