English
Related papers

Related papers: Online Controlled Experiments for Personalised e-C…

200 papers

Using offline datasets to evaluate conversational agents often fails to cover rare scenarios or to support testing new policies. This has motivated the use of controllable user simulators for targeted, counterfactual evaluation, typically…

Artificial Intelligence · Computer Science 2026-05-13 Guy Tennenholtz , Ofer Meshi , Amir Globerson , Uri Shalit , Jihwan Jeong , Craig Boutilier

A/B testing is a widely-used paradigm within marketing optimization because it promises identification of causal effects and because it is implemented out of the box in most messaging delivery software platforms. Modern businesses, however,…

Machine Learning · Computer Science 2023-05-03 Schaun Wheeler

Participants in online experiments often enroll over time, which can compromise sample representativeness due to temporal shifts in covariates. This issue is particularly critical in A/B tests, online controlled experiments extensively used…

General Economics · Economics 2026-03-30 Chen Wang , Shichao Han , Shan Huang

A/B experiments are commonly used in research to compare the effects of changing one or more variables in two different experimental groups - a control group and a treatment group. While the benefits of using A/B experiments are widely…

Software Engineering · Computer Science 2023-09-26 Andrew Hornback , Sungeun An , Scott Bunin , Stephen Buckley , John Kos , Ashok Goel

Latest research revealed a considerable lack of reliability within user feedback and discussed striking impacts for the assessment of adaptive web systems and content personalisation approaches, e.g. ranking errors, systematic biases to…

Human-Computer Interaction · Computer Science 2018-02-19 Kevin Jasberg , Sergej Sizov

Recommender systems are widely used AI applications designed to help users efficiently discover relevant items. The effectiveness of such systems is tied to the satisfaction of both users and providers. However, user satisfaction is complex…

Information Retrieval · Computer Science 2024-11-05 Ali Elahi , Armin Zirak

Offline evaluations of recommender systems attempt to estimate users' satisfaction with recommendations using static data from prior user interactions. These evaluations provide researchers and developers with first approximations of the…

Information Retrieval · Computer Science 2020-01-28 Mucun Tian , Michael D. Ekstrand

A/B testing is gaining attention in the automotive sector as a promising tool to measure causal effects from software changes. Different from the web-facing businesses, where A/B testing has been well-established, the automotive domain…

Software Engineering · Computer Science 2021-11-12 Yuchu Liu , David Issa Mattos , Jan Bosch , Helena Holmström Olsson , Jonn Lantz

Recommender systems operate in an inherently dynamical setting. Past recommendations influence future behavior, including which data points are observed and how user preferences change. However, experimenting in production systems with real…

Information Retrieval · Computer Science 2020-11-17 Karl Krauth , Sarah Dean , Alex Zhao , Wenshuo Guo , Mihaela Curmei , Benjamin Recht , Michael I. Jordan

Randomized Controlled Trials (RCTs), or A/B testing, have become the gold standard for optimizing various operational policies on online platforms. However, RCTs on these platforms typically cover a limited number of discrete treatment…

Econometrics · Economics 2026-02-06 Zhiqi Zhang , Zhiyu Zeng , Ruohan Zhan , Dennis Zhang

Personalized recommendations form an important part of today's internet ecosystem, helping artists and creators to reach interested users, and helping users to discover new and engaging content. However, many users today are skeptical of…

Cryptography and Security · Computer Science 2024-01-09 Allegra Laro , Yanqing Chen , Hao He , Babak Aghazadeh

The reliability of controlled experiments, commonly referred to as "A/B tests," is often compromised by network interference, where the outcomes of individual units are influenced by interactions with others. Significant challenges in this…

Machine Learning · Statistics 2024-07-02 Yuan Yuan , Kristen M. Altenburger

Online field experiments are the gold-standard way of evaluating changes to real-world interactive machine learning systems. Yet our ability to explore complex, multi-dimensional policy spaces - such as those found in recommendation and…

Machine Learning · Statistics 2019-04-30 Benjamin Letham , Eytan Bakshy

Experimentation is widely utilized for causal inference and data-driven decision-making across disciplines. In an A/B experiment, for example, an online business randomizes two different treatments (e.g., website designs) to their customers…

Methodology · Statistics 2025-01-15 Wenxuan Guo , JungHo Lee , Panos Toulis

Artificial Intelligence is being employed by humans to collaboratively solve complicated tasks for search and rescue, manufacturing, etc. Efficient teamwork can be achieved by understanding user preferences and recommending different…

Information Retrieval · Computer Science 2023-01-20 Lakshita Dodeja , Pradyumna Tambwekar , Erin Hedlund-Botti , Matthew Gombolay

Most modern recommendation algorithms are data-driven: they generate personalized recommendations by observing users' past behaviors. A common assumption in recommendation is that how a user interacts with a piece of content (e.g., whether…

Computers and Society · Computer Science 2024-05-12 Sarah H. Cen , Andrew Ilyas , Jennifer Allen , Hannah Li , Aleksander Madry

A/B testing remains the gold standard for evaluating e-commerce UI changes, yet it diverts traffic, takes weeks to reach significance, and risks harming user experience. We introduce SimGym, a scalable system for rapid offline A/B testing…

Recommender systems have generated tremendous value for both users and businesses, drawing significant attention from academia and industry alike. However, due to practical constraints, academic research remains largely confined to offline…

Information Retrieval · Computer Science 2025-09-09 Kuan Zou , Aixin Sun

We describe our framework, deployed at Facebook, that accounts for interference between experimental units through cluster-randomized experiments. We document this system, including the design and estimation procedures, and detail insights…

Social and Information Networks · Computer Science 2020-12-17 Brian Karrer , Liang Shi , Monica Bhole , Matt Goldman , Tyrone Palmer , Charlie Gelman , Mikael Konutgan , Feng Sun

We consider situations where consumers are aware that a statistical model determines the price of a product based on their observed behavior. Using a novel experiment varying the context similarity between participant data and a product, we…

General Economics · Economics 2024-11-14 Inácio Bó , Li Chen , Rustamdjan Hakimov