中文
相关论文

相关论文: Risk-aware product decisions in A/B tests with mul…

200 篇论文

We propose a multi-metric flexible Bayesian framework to support efficient interim decision-making in multi-arm multi-stage phase II clinical trials. Multi-arm multi-stage phase II studies increase the efficiency of drug development, but…

应用统计 · 统计学 2023-12-15 Suzanne M. Dufault , Angela M. Crook , Katie Rolfe , Patrick P. J. Phillips

In multiple testing scenarios, typically the sign of a parameter is inferred when its estimate exceeds some significance threshold in absolute value. Typically, the significance threshold is chosen to control the experimentwise type I error…

统计方法学 · 统计学 2018-01-03 Chaoyu Yu , Peter D. Hoff

A/B testing, or online experiment is a standard business strategy to compare a new product with an old one in pharmaceutical, technological, and traditional industries. Major challenges arise in online experiments of two-sided marketplace…

机器学习 · 计算机科学 2022-11-04 Chengchun Shi , Xiaoyu Wang , Shikai Luo , Hongtu Zhu , Jieping Ye , Rui Song

Consider a scenario where a harmfulness evaluation metric intended to filter unsafe responses from a Large Language Model. When applied to individual harmful prompt-response pairs, it correctly flags them as unsafe by assigning a high-risk…

人工智能 · 计算机科学 2025-02-14 Ora Nova Fandina , Leshem Choshen , Eitan Farchi , George Kour , Yotam Perlitz , Orna Raz

We propose a distributional framework for benchmarking socio-technical risks of foundation models with quantified statistical significance. Our approach hinges on a new statistical relative testing based on first and second order stochastic…

Metrics can be used by businesses to make more objective decisions based on data. Software startups in particular are characterized by the uncertain or even chaotic nature of the contexts in which they operate. Using data in the form of…

综合文献 · 计算机科学 2019-01-16 Kai-Kristian Kemell , Xiaofeng Wang , Anh Nguyen-Duc , Jason Grendus , Tuure Tuunanen , Pekka Abrahamsson

User-randomized A/B testing has emerged as the gold standard for online experimentation. However, when this kind of approach is not feasible due to legal, ethical or practical considerations, experimenters have to consider alternatives like…

统计方法学 · 统计学 2025-06-17 Paul Missault , Lorenzo Masoero , Christian Delbé , Thomas Richardson , Guido Imbens

Quality control is a crucial activity performed by manufacturing enterprises to ensure that their products meet quality standards and avoid potential damage to the brand's reputation. The decreased cost of sensors and connectivity enabled…

机器学习 · 计算机科学 2022-11-28 Jože M. Rožanec , Luka Bizjak , Elena Trajkova , Patrik Zajec , Jelle Keizer , Blaž Fortuna , Dunja Mladenić

Software companies have widely used online A/B testing to evaluate the impact of a new technology by offering it to groups of users and comparing it against the unmodified product. However, running online A/B testing needs not only efforts…

软件工程 · 计算机科学 2024-08-12 Jie JW Wu

Advanced AI technologies are serving humankind in a number of ways, from healthcare to manufacturing. Advanced automated machines are quite expensive, but the end output is supposed to be of the highest possible quality. Depending on the…

软件工程 · 计算机科学 2022-10-03 Gokul Yenduri , Thippa Reddy Gadekallu

We address the problem of A/B testing, a widely used protocol for evaluating the potential improvement achieved by a new decision system compared to a baseline. This protocol segments the population into two subgroups, each exposed to a…

机器学习 · 统计学 2025-06-16 Otmane Sakhi , Alexandre Gilotte , David Rohde

Applying Artificial Intelligence (AI) and Machine Learning (ML) in critical contexts, such as medicine, requires the implementation of safety measures to reduce risks of harm in case of prediction errors. Spotting ML failures is of…

Empirical research in the social and medical sciences frequently involves testing multiple hypotheses simultaneously, increasing the risk of false positives due to chance. Classical multiple testing procedures, such as the Bonferroni…

计量经济学 · 经济学 2025-07-29 Sebastian Calonico , Sebastian Galiani

The problem of simultaneously testing the marginal distributions of sequentially monitored, independent data streams is considered. The decisions for the various testing problems can be made at different times, using data from all streams,…

统计方法学 · 统计学 2023-04-21 Yiming Xing , Georgios Fellouris

In this paper we consider the problem of multiple testing when the hypotheses are dependent. In most of the existing literature, either Bayesian or non-Bayesian, the decision rules mainly focus on the validity of the test procedure rather…

统计方法学 · 统计学 2018-07-10 Noirrit K. Chandra , Sourabh Bhattacharya

Designing products to meet consumers' preferences is essential for a business's success. We propose the Gradient-based Survey (GBS), a discrete choice experiment for multiattribute product design. The experiment elicits consumer preferences…

机器学习 · 统计学 2023-10-19 Mingzhang Yin , Ruijiang Gao , Weiran Lin , Steven M. Shugan

Embedded systems interaction with environment inherently complicates understanding of requirements and their correct implementation. However, product uncertainty is highest during early stages of development. Design verification is an…

软件工程 · 计算机科学 2011-12-20 Hara Gopal Mani Pakala , Dr. Plh Varaprasad , Dr. Raju Kvsvn , Dr. Ibrahim Khan

Recommender systems have become an integral part of online platforms, providing personalized recommendations for purchases, content consumption, and interpersonal connections. These systems consist of two sides: the producer side comprises…

统计方法学 · 统计学 2023-11-08 Yan Wang , Shan Ba

When developing a new networking algorithm, it is established practice to run a randomized experiment, or A/B test, to evaluate its performance. In an A/B test, traffic is randomly allocated between a treatment group, which uses the new…

网络与互联网体系结构 · 计算机科学 2021-10-04 Bruce Spang , Veronica Hannan , Shravya Kunamalla , Te-Yuan Huang , Nick McKeown , Ramesh Johari

Technical and legal debates frequently suggest that "accuracy" is an objective, measurable, and purely technical property. We challenge this view, showing that evaluating AI performance fundamentally depends on context-dependent normative…