中文
相关论文

相关论文: Evaluating Decision Rules Across Many Weak Experim…

200 篇论文

A/B testing plays a central role in data-driven product development, guiding launch decisions for new features and designs. However, treatment effect estimates are often noisy due to short horizons, early stopping, and slowly accumulating…

统计方法学 · 统计学 2025-11-27 Xinran Li

Controlled experimentation, also called A/B testing, is widely adopted to accelerate product innovations in the online world. However, how fast we innovate can be limited by how we run experiments. Most experiments go through a "ramp up"…

应用统计 · 统计学 2018-01-26 Ya Xu , Weitao Duan , Shaochen Huang

Committee-selection problems arise in many contexts and applications, and there has been increasing interest within the social choice research community on identifying which properties are satisfied by different multi-winner voting rules.…

人工智能 · 计算机科学 2025-08-11 Joshua Caiata , Ben Armstrong , Kate Larson

Innovations across science and industry are evaluated using randomized trials (a.k.a. A/B tests). While simple and robust, such static designs are inefficient or infeasible for testing many hypotheses. Adaptive designs can greatly improve…

机器学习 · 计算机科学 2024-08-09 Jimmy Wang , Ethan Che , Daniel R. Jiang , Hongseok Namkoong

A/B testing is an important decision making tool in product development because can provide an accurate estimate of the average treatment effect of a new features, which allows developers to understand how the business impact of new changes…

应用统计 · 统计学 2019-03-22 Guillaume Saint-Jacques , James Eric Sorenson , Nanyu Chen , Ya Xu

A/B tests, also known as randomized controlled experiments (RCTs), are the gold standard for evaluating the impact of new policies, products, or decisions. However, these tests can be costly in terms of time and resources, potentially…

机器学习 · 统计学 2025-01-03 Shima Nassiri , Mohsen Bayati , Joe Cooprider

Online experiments %in which experimental units receive a sequence of treatments over time are frequently employed in many technological companies to evaluate the performance of a newly developed policy, product, or treatment relative to a…

计量经济学 · 经济学 2025-01-14 Ke Sun , Linglong Kong , Hongtu Zhu , Chengchun Shi

Randomized trials, also known as A/B tests, are used to select between two policies: a control and a treatment. Given a corresponding set of features, we can ideally learn an optimized policy P that maps the A/B test data features to action…

机器学习 · 计算机科学 2018-06-08 Elon Portugaly , Joseph J. Pfeiffer

The effectiveness of recommendation systems is pivotal to user engagement and satisfaction in online platforms. As these recommendation systems increasingly influence user choices, their evaluation transcends mere technical performance and…

信息检索 · 计算机科学 2024-01-15 Aryan Jadon , Avinash Patil

When users rate objects, a sophisticated algorithm that takes into account ability or reputation may produce a fairer or more accurate aggregation of ratings than the straightforward arithmetic average. Recently a number of authors have…

信息检索 · 计算机科学 2015-03-13 Matus Medo , Joseph Rushton Wakeling

There has long been debates on how we could interpret neural networks and understand the decisions our models make. Specifically, why deep neural networks tend to be error-prone when dealing with samples that output low softmax scores. We…

计算机视觉与模式识别 · 计算机科学 2018-12-04 Simiao Zuo , Jialin Wu

Motivated by the psychological literature on the "peak-end rule" for remembered experience, we perform an analysis within a random walk framework of a discrete choice model where agents' future choices depend on the peak memory of their…

统计力学 · 物理学 2015-06-01 Rosemary J. Harris

The goal of a next basket recommendation (NBR) system is to recommend items for the next basket for a user, based on the sequence of their prior baskets. Recently, a number of methods with complex modules have been proposed that claim…

信息检索 · 计算机科学 2023-03-10 Ming Li , Sami Jullien , Mozhdeh Ariannezhad , Maarten de Rijke

On e-commerce platforms, predicting if two products are compatible with each other is an important functionality to achieve trustworthy product recommendation and search experience for consumers. However, accurately predicting product…

机器学习 · 计算机科学 2022-06-29 Rongzhi Zhang , Rebecca West , Xiquan Cui , Chao Zhang

Current approaches to A/B testing in networks focus on limiting interference, the concern that treatment effects can "spill over" from treatment nodes to control nodes and lead to biased causal effect estimation. Prominent methods for…

机器学习 · 计算机科学 2020-04-16 Zahra Fatemi , Elena Zheleva

A major challenge in data-driven decision-making is accurate policy evaluation-i.e., guaranteeing that a learned decision-making policy achieves the promised benefits. A popular strategy is model-based policy evaluation, which estimates a…

机器学习 · 统计学 2026-02-10 Hamsa Bastani , Osbert Bastani , Bryce McLaughlin

Design of experiments and estimation of treatment effects in large-scale networks, in the presence of strong interference, is a challenging and important problem. Most existing methods' performance deteriorates as the density of the network…

统计方法学 · 统计学 2020-12-15 Preetam Nandy , Kinjal Basu , Shaunak Chatterjee , Ye Tu

The robot learning community has made great strides in recent years, proposing new architectures and showcasing impressive new capabilities; however, the dominant metric used in the literature, especially for physical experiments, is…

Recommender systems have become crucial in the modern digital landscape, where personalized content, products, and services are essential for enhancing user experience. This paper explores statistical models for recommender systems,…

统计方法学 · 统计学 2024-08-13 Disha Ghandwani , Trevor Hastie

Randomized experiments, or A/B testing, are the gold standard for evaluating interventions, yet they remain underutilized in inventory management. This study addresses this gap by analyzing A/B testing strategies in multi-item, multi-period…

统计方法学 · 统计学 2026-02-03 Xinqi Chen , Xingyu Bai , Zeyu Zheng , Nian Si