English
Related papers

Related papers: Measuring Diagnostic Test Performance Using Imperf…

200 papers

From scientific experiments to online A/B testing, the previously observed data often affects how future experiments are performed, which in turn affects which data will be collected. Such adaptivity introduces complex correlations between…

Machine Learning · Statistics 2018-01-03 Xinkun Nie , Xiaoying Tian , Jonathan Taylor , James Zou

The repeated community-wide reuse of test sets in popular benchmark problems raises doubts about the credibility of reported test-error rates. Verifying whether a learned model is overfitted to a test set is challenging as independent test…

Machine Learning · Computer Science 2019-11-15 Roman Werpachowski , András György , Csaba Szepesvári

The purpose of this project was to collect and analyse data about the comparability and real-life applicability of published results focusing on Microsoft Windows malware, more specifically the impact of dataset size and testing dataset…

Cryptography and Security · Computer Science 2022-06-14 David Illes

The proximal causal inference framework enables the identification and estimation of causal effects in the presence of unmeasured confounding by leveraging two disjoint sets of observed strong proxies: negative control treatments and…

Methodology · Statistics 2025-12-16 Antonio Olivas-Martinez , Peter B. Gilbert , Andrea Rotnitzky

Anomaly detection methods have demonstrated remarkable success across various applications. However, assessing their performance, particularly at the pixel-level, presents a complex challenge due to the severe imbalance that is most…

Computer Vision and Pattern Recognition · Computer Science 2023-10-26 Mehdi Rafiei , Toby P. Breckon , Alexandros Iosifidis

An important issue in medical image processing is to be able to estimate not only the performances of algorithms but also the precision of the estimation of these performances. Reporting precision typically amounts to reporting…

Computer Vision and Pattern Recognition · Computer Science 2023-05-25 Rosana El Jurdi , Olivier Colliot

A goodness-of-fit test for the fitting of a parametric model to data obtained from a detector with finite resolution and limited acceptance is proposed. The parameters of the model are found by minimization of a statistic that is used for…

Data Analysis, Statistics and Probability · Physics 2015-03-17 N. D. Gagunashvili

Pragmatic trials increasingly define outcomes using real-world data such as electronic health records, where assessments are collected during routine care rather than at fixed timepoints. Consequently, these uncontrolled assessments may be…

The partial conjunction null hypothesis is tested in order to discover a signal that is present in multiple studies. The standard approach of carrying out a multiple test procedure on the partial conjunction (PC) $p$-values can be extremely…

Methodology · Statistics 2024-06-14 Thorsten Dickhaus , Ruth Heller , Anh-Tuan Hoang , Yosef Rinott

Robustness checks are routine in empirical work, but there is no standard statistical procedure to formally measure what one can learn from them. I propose a "robustness radius" measure to quantify the amount by which the robustness checks…

Econometrics · Economics 2026-02-24 Brenda Prallon

The policy relevant treatment effect (PRTE) measures the average effect of switching from a status-quo policy to a counterfactual policy. Estimation of the PRTE involves estimation of multiple preliminary parameters, including propensity…

Econometrics · Economics 2020-07-17 Yuya Sasaki , Takuya Ura

Multiple hypothesis testing is a fundamental problem in high dimensional inference, with wide applications in many scientific fields. In genome-wide association studies, tens of thousands of tests are performed simultaneously to find if any…

Methodology · Statistics 2011-11-16 Jianqing Fan , Xu Han , Weijie Gu

In epidemic or pandemic situations, resources for testing the infection status of individuals may be scarce. Although group testing can help to significantly increase testing capabilities, the (repeated) testing of entire populations can…

Populations and Evolution · Quantitative Biology 2021-10-29 Günther Koliander , Georg Pichler

We present an extension to the robust phase estimation protocol, which can identify incorrect results that would otherwise lie outside the expected statistical range. Robust phase estimation is increasingly a method of choice for…

Sequential tests and their implied confidence sequences, which are valid at arbitrary stopping times, promise flexible statistical inference and on-the-fly decision making. However, strong guarantees are limited to parametric sequential…

Methodology · Statistics 2024-03-12 Aurelien Bibaut , Nathan Kallus , Michael Lindon

Across domains such as medicine, employment, and criminal justice, predictive models often target labels that imperfectly reflect the outcomes of interest to experts and policymakers. For example, clinical risk assessments deployed to…

Machine Learning · Computer Science 2023-05-19 Luke Guerdan , Amanda Coston , Kenneth Holstein , Zhiwei Steven Wu

Statistical hypothesis tests typically use prespecified sample sizes, yet data often arrive sequentially. Interim analyses invalidate classical error guarantees, while existing sequential methods require rigid testing preschedules or incur…

Methodology · Statistics 2026-02-17 Chris Holmes , Stephen Walker

We apply multiple testing procedures to the validation of estimated default probabilities in credit rating systems. The goal is to identify rating classes for which the probability of default is estimated inaccurately, while still…

Applications · Statistics 2010-06-28 Sebastian Döhler

The positive false discovery rate (pFDR) is a useful overall measure of errors for multiple hypothesis testing, especially when the underlying goal is to attain one or more discoveries. Control of pFDR critically depends on how much…

Statistics Theory · Mathematics 2011-11-09 Zhiyi Chi

Anomaly-based intrusion detection promises to detect novel or unknown attacks on industrial control systems by modeling expected system behavior and raising corresponding alarms for any deviations.As manually creating these behavioral…

Cryptography and Security · Computer Science 2022-05-20 Dominik Kus , Eric Wagner , Jan Pennekamp , Konrad Wolsing , Ina Berenice Fink , Markus Dahlmanns , Klaus Wehrle , Martin Henze