English
Related papers

Related papers: Risk-aware product decisions in A/B tests with mul…

200 papers

A/B testing refers to the statistical procedure of conducting an experiment to compare two treatments, A and B, applied to different testing subjects. It is widely used by technology companies such as Facebook, LinkedIn, and Netflix, to…

Methodology · Statistics 2026-05-12 Victoria Pokhiko , Qiong Zhang , Lulu Kang , D'arcy P. Mays

Time series anomaly detection is widely used in IoT and cyber-physical systems, yet its evaluation remains challenging due to diverse application objectives and heterogeneous metric assumptions. This study introduces a problem-oriented…

Artificial Intelligence · Computer Science 2026-05-15 Kaixiang Yang , Jiarong Liu , Yupeng Song , Shuanghua Yang , Yujue Zhou

Context: Software testability is the degree to which a software system or a unit under test supports its own testing. To predict and improve software testability, a large number of techniques and metrics have been proposed by both…

Software Engineering · Computer Science 2018-12-07 Vahid Garousi , Michael Felderer , Feyza Nur Kilicaslan

In safety-critical applications a probabilistic model is usually required to be calibrated, i.e., to capture the uncertainty of its predictions accurately. In multi-class classification, calibration of the most confident predictions only is…

Machine Learning · Statistics 2022-09-30 David Widmann , Fredrik Lindsten , Dave Zachariah

In many Deep Reinforcement Learning (RL) problems, decisions in a trained policy vary in significance for the expected safety and performance of the policy. Since RL policies are very complex, testing efforts should concentrate on states in…

Machine Learning · Computer Science 2024-11-13 Stefan Pranger , Hana Chockler , Martin Tappler , Bettina Könighofer

Contemporary sample size calculations for external validation of risk prediction models require users to specify fixed values of assumed model performance metrics alongside target precision levels (e.g., 95% CI widths). However, due to the…

Applications · Statistics 2026-02-13 Mohsen Sadatsafavi , Paul Gustafson , Solmaz Setayeshgar , Laure Wynants , Richard D Riley

You measure the value of a quantity x for a number of systems (cells, molecules, people, chunks of metal, DNA vectors, etc.). You repeat the whole set of measures in different occasions or assays, which you try to design as equal to one…

Quantitative Methods · Quantitative Biology 2013-11-04 Pablo Echenique-Robba , María Alejandra Nelo-Bazán , José A. Carrodeguas

Safety evaluation of self-driving technologies has been extensively studied. One recent approach uses Monte Carlo based evaluation to estimate the occurrence probabilities of safety-critical events as safety measures. These Monte Carlo…

Methodology · Statistics 2019-07-19 Zhiyuan Huang , Mansur Arief , Henry Lam , Ding Zhao

The first part of this thesis focuses on maximizing the overall recommendation accuracy. This accuracy is usually evaluated with some user-oriented metric tailored to the recommendation scenario, but because recommendation is usually…

Information Retrieval · Computer Science 2023-11-14 Roger Zhe Li

In a well-calibrated risk prediction model, the average predicted probability is close to the true event rate for any given subgroup. Such models are reliable across heterogeneous populations and satisfy strong notions of algorithmic…

Machine Learning · Computer Science 2023-07-31 Jean Feng , Alexej Gossmann , Romain Pirracchio , Nicholas Petrick , Gene Pennello , Berkman Sahiner

While there exists a large amount of literature on the general challenges of and best practices for trustworthy online A/B testing, there are limited studies on sample size estimation, which plays a crucial role in trustworthy and efficient…

Methodology · Statistics 2023-08-21 Jing Zhou , Jiannan Lu , Anas Shallah

Selective Classification, wherein models can reject low-confidence predictions, promises reliable translation of machine-learning based classification systems to real-world scenarios such as clinical diagnostics. While current evaluation of…

In a world of utility-driven marketing, each company acts as an adversary to other contenders, with all having competing interests. A major challenge for companies launching a new product is that, despite testing, flaws in their product can…

Applications · Statistics 2025-06-02 Pablo G. Arce , Sonali Das , David Ríos Insua

The rise of biomedical foundation models creates new hurdles in model testing and authorization, given their broad capabilities and susceptibility to complex distribution shifts. We suggest tailoring robustness tests according to…

Software Engineering · Computer Science 2025-09-01 R. Patrick Xian , Noah R. Baker , Tom David , Qiming Cui , A. Jay Holmgren , Stefan Bauer , Madhumita Sushil , Reza Abbasi-Asl

The two-trials rule for drug approval requires "at least two adequate and well-controlled studies, each convincing on its own, to establish effectiveness". This is usually implemented by requiring two significant pivotal trials and is the…

Methodology · Statistics 2023-11-10 Leonhard Held

Ordinal user-provided ratings across multiple items are frequently encountered in both scientific and commercial applications. Whilst recommender systems are known to do well on these type of data from a predictive point of view, their…

Methodology · Statistics 2025-03-05 Sjoerd Hermes

A growing demand for handling uncertainties and risks in performance-driven building design decision-making has challenged conventional design methods. Thus, researchers in this field lean towards viable alternatives to using deterministic…

Numerical Analysis · Mathematics 2021-09-20 Fatemeh Shahsavari , Jeffrey D. Hart , Wei Yan

A well-known approach for identifying defect-prone parts of software in order to focus testing is to use different kinds of product metrics such as size or complexity. Although this approach has been evaluated in many contexts, the question…

Software Engineering · Computer Science 2014-02-05 Frank Elberzhager , Stephan Kremer , Jürgen Münch , Danilo Assmann

Modern high-throughput biomedical devices routinely produce data on a large scale, and the analysis of high-dimensional datasets has become commonplace in biomedical studies. However, given thousands or tens of thousands of measured…

Methodology · Statistics 2022-02-28 Vladimir Vutov , Thorsten Dickhaus

Applying software defect esimation techniques and presenting this information in a compact and impactful decision table can clearly illustrate to collaborative groups how critical this position is in the overall development cycle. The Test…

Software Engineering · Computer Science 2007-11-13 James Cusick