English
Related papers

Related papers: Evaluating Decision Rules Across Many Weak Experim…

200 papers

Randomised Controlled Trials (RCTs) are the gold standard for estimating treatment effects across many fields of science. Technology companies have adopted A/B-testing methods as a modern RCT counterpart, where end-users are randomly…

Social and Information Networks · Computer Science 2024-09-20 Olivier Jeunen

I examine a conceptual model of a recommendation system (RS) with user inflow and churn dynamics. When inflow and churn balance out, the user distribution reaches a steady state. Changing the recommendation algorithm alters the steady state…

Information Retrieval · Computer Science 2024-10-31 Shichao Ma

The use of relevant metrics of software systems could improve various software engineering tasks, but identifying relationships among metrics is not simple and can be very time consuming. Recommender systems can help with this…

Software Engineering · Computer Science 2018-01-23 Maral Azizi , Hyunsook Do

The ongoing rapid development of the e-commercial and interest-base websites make it more pressing to evaluate objects' accurate quality before recommendation by employing an effective reputation system. The objects' quality are often…

Physics and Society · Physics 2018-07-23 Leilei Wu , Zhuoming Ren , Xiao-Long Ren , Jianlin Zhang , Linyuan Lü

Randomized experiments is a key part of product development in the tech industry. It is often necessary to run programs of exclusive experiments, i.e., experiments that cannot be run on the same units during the same time. These programs…

Methodology · Statistics 2020-12-21 Mårten Schultzberg , Oskar Kjellin , Johan Rydberg

In the criminal legal context, risk assessment algorithms are touted as data-driven, well-tested tools. Studies known as validation tests are typically cited by practitioners to show that a particular risk assessment algorithm has…

Computers and Society · Computer Science 2020-12-01 Benjamin Laufer

Adaptive experiments, including efficient average treatment effect estimation and multi-armed bandit algorithms, have garnered attention in various applications, such as social experiments, clinical trials, and online advertisement…

Methodology · Statistics 2021-03-24 Masahiro Kato

Online decision making aims to learn the optimal decision rule by making personalized decisions and updating the decision rule recursively. It has become easier than before with the help of big data, but new challenges also come along.…

Machine Learning · Statistics 2020-10-16 Haoyu Chen , Wenbin Lu , Rui Song

A/B tests serve the purpose of reliably identifying the effect of changes introduced in online services. It is common for online platforms to run a large number of simultaneous experiments by splitting incoming user traffic randomly in…

Machine Learning · Computer Science 2022-10-18 Alexander Buchholz , Vito Bellini , Giuseppe Di Benedetto , Yannik Stein , Matteo Ruffini , Fabian Moerchen

Rules based approaches for data quality solutions often use business rules or integrity rules for data monitoring purpose. Integrity rules are constraints on data derived from business rules into a formal form in order to allow…

Software Engineering · Computer Science 2017-04-21 Thanh Thoa Pham Thi , Markus Helfert

This article studies the benefits of using spatially randomized experimental designs which partition the experimental area into distinct, non-overlapping units with treatments assigned randomly. Such designs offer improved policy evaluation…

Statistics Theory · Mathematics 2025-11-18 Ying Yang , Chengchun Shi , Fang Yao , Shouyang Wang , Hongtu Zhu

A/B experiments are commonly used in research to compare the effects of changing one or more variables in two different experimental groups - a control group and a treatment group. While the benefits of using A/B experiments are widely…

Software Engineering · Computer Science 2023-09-26 Andrew Hornback , Sungeun An , Scott Bunin , Stephen Buckley , John Kos , Ashok Goel

Reducing negative user experiences is essential for the success of recommendation platforms. Exposing users to inappropriate content could not only adversely affect users' psychological well-beings, but also potentially drive users away…

Information Retrieval · Computer Science 2025-02-18 Chenghui Yu , Peiyi Li , Haoze Wu , Yiri Wen , Bingfeng Deng , Hongyu Xiong

Imagine a food recommender system -- how would we check if it is \emph{causing} and fostering unhealthy eating habits or merely reflecting users' interests? How much of a user's experience over time with a recommender is caused by the…

Machine Learning · Computer Science 2021-01-13 Sirui Yao , Yoni Halpern , Nithum Thain , Xuezhi Wang , Kang Lee , Flavien Prost , Ed H. Chi , Jilin Chen , Alex Beutel

A new Thurstonian rating scale model uses a variable decision rule (VDR) that incorporates three previously formulated, distinct decision rules. The model includes probabilities for choosing each rule, along with Gaussian representation and…

Methodology · Statistics 2016-02-01 Burton Rosner , Greg Kochanski

Considerable interest has recently been focused on studying multiple phenotypes simultaneously in both epidemiological and genomic studies, either to capture the multidimensionality of complex disorders or to understand shared etiology of…

Methodology · Statistics 2015-11-26 Denis Agniel , Katherine P. Liao , Tianxi Cai

A/B testing is the foundation of decision-making in online platforms, yet social products often suffer from network interference: user interactions cause treatment effects to spill over into the control group. Such spillovers bias causal…

Social and Information Networks · Computer Science 2026-02-10 Xu Min , Zhaoxu Yang , Kaixuan Tan , Juan Yan , Xunbin Xiong , Zihao Zhu , Kaiyu Zhu , Fenglin Cui , Yang Yang , Sihua Yang , Jianhui Bu

Netflix is an internet entertainment service that routinely employs experimentation to guide strategy around product innovations. As Netflix grew, it had the opportunity to explore increasingly specialized improvements to its service, which…

Risk scoring systems are widely used in high-stakes domains to assist decision-making. However, existing approaches often focus on optimizing predictive accuracy or likelihood-based criteria, which may not align with the main goal of…

Machine Learning · Computer Science 2026-04-07 Wenhao Chi , Ş. İlker Birbil

Modern online experimentation faces two bottlenecks: scarce traffic forces tough choices on which variants to test, and post-hoc insight extraction is manual, inconsistent, and often content-agnostic. Meanwhile, organizations underuse…

Artificial Intelligence · Computer Science 2026-02-17 Zhengmian Hu , Lei Shi , Ritwik Sinha , Justin Grover , David Arbour