English
Related papers

Related papers: A statistical framework for planning and analysing…

200 papers

Lack of reliability is a well-known issue for reinforcement learning (RL) algorithms. This problem has gained increasing attention in recent years, and efforts to improve it have grown substantially. To aid RL researchers and production…

Machine Learning · Statistics 2020-02-14 Stephanie C. Y. Chan , Samuel Fishman , John Canny , Anoop Korattikara , Sergio Guadarrama

Repeated measures of biomarkers have the potential of explaining hazards of survival outcomes. In practice, these measurements are intermittently measured and are known to be subject to substantial measurement error. Joint modelling of…

Applications · Statistics 2019-12-12 Lisa McFetridge , Ozgur Asar , Jonas Wallin

In this paper the accuracy and robustness of quality measures for the assessment of machine learning models are investigated. The prediction quality of a machine learning model is evaluated model-independent based on a cross-validation…

Machine Learning · Statistics 2024-10-07 Thomas Most , Lars Gräning , Sebastian Wolff

Large language models (LLMs) are stochastic, and not all models give deterministic answers, even when setting temperature to zero with a fixed random seed. However, few benchmark studies attempt to quantify uncertainty, partly due to the…

Computation and Language · Computer Science 2025-06-30 Robert E. Blackwell , Jon Barry , Anthony G. Cohn

Participant level meta-analysis across multiple studies increases the sample size for pooled analyses, thereby improving precision in effect estimates and enabling subgroup analyses. For analyses involving biomarker measurements as an…

Reliability is an essential measure of how closely observed scores represent latent scores (reflecting constructs), assuming some latent variable measurement model. We present a general theoretical framework of reliability, placing emphasis…

Methodology · Statistics 2024-10-29 Yang Liu , Jolynn Pek , Alberto Maydeu-Olivares

The advent of modern data collection and processing techniques has seen the size, scale, and complexity of data grow exponentially. A seminal step in leveraging these rich datasets for downstream inference is understanding the…

Applications · Statistics 2024-07-30 Zeyi Wang , Eric Bridgeford , Shangsi Wang , Joshua T. Vogelstein , Brian Caffo

To develop rigorous knowledge about ML models -- and the systems in which they are embedded -- we need reliable measurements. But reliable measurement is fundamentally challenging, and touches on issues of reproducibility, scalability,…

Machine Learning · Computer Science 2024-08-13 A. Feder Cooper

The assessment of imaging biomarkers is critical for advancing precision medicine and improving disease characterization. Despite the availability of methods to derive disease heterogeneity metrics in imaging studies, a robust framework for…

Repeating an imperfect biomarker test based on an initial result can introduce bias and influence misclassification risk. For example, in some blood donation settings, blood donors' hemoglobin is remeasured when the initial measurement…

Applications · Statistics 2026-02-17 Supun Manathunga , Mart P. Janssen , Yu Luo , W. Alton Russell , Mart Pothast

Ratio-based biomarkers (RBBs), such as the proportion of necrotic tissue within a tumor, are widely used in clinical practice to support diagnosis, prognosis, and treatment planning. These biomarkers are typically estimated from…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Jiameng Li , Teodora Popordanoska , Aleksei Tiulpin , Sebastian G. Gruber , Frederik Maes , Matthew B. Blaschko

Randomized benchmarking (RB) is an efficient and robust method to characterize gate errors in quantum circuits. Averaging over random sequences of gates leads to estimates of gate errors in terms of the average fidelity. These estimates are…

Quantum Physics · Physics 2019-09-11 Jonas Helsen , Joel J. Wallman , Steven T. Flammia , Stephanie Wehner

Systems that answer questions by reviewing the scientific literature are becoming increasingly feasible. To draw reliable conclusions, these systems should take into account the quality of available evidence from different studies, placing…

Computation and Language · Computer Science 2025-09-23 Jianyou Wang , Weili Cao , Longtian Bao , Youze Zheng , Gil Pasternak , Kaicheng Wang , Xiaoyue Wang , Ramamohan Paturi , Leon Bergen

Benchmarking is crucial for testing and validating any system, even more so in real-time systems. Typical real-time applications adhere to well-understood abstractions: they exhibit a periodic behavior, operate on a well-defined working…

Software Engineering · Computer Science 2022-08-02 Mattia Nicolella , Shahin Roozkhosh , Denis Hoornaert , Andrea Bastoni , Renato Mancuso

Replicability and reproducibility of experimental results are primary concerns in all the areas of science and IR is not an exception. Besides the problem of moving the field towards more reproducible experimental practices and protocols,…

Information Retrieval · Computer Science 2020-10-27 Timo Breuer , Nicola Ferro , Norbert Fuhr , Maria Maistro , Tetsuya Sakai , Philipp Schaer , Ian Soboroff

For a general standardized testing algorithm designed to evaluate a specific aspect of a robot's performance, several key expectations are commonly imposed. Beyond accuracy (i.e., closeness to a typically unknown ground-truth reference) and…

Robotics · Computer Science 2025-12-22 Bowen Weng , Linda Capito , Guillermo A. Castillo , Dylan Khor

The ability to replicate predictions by machine learning (ML) or artificial intelligence (AI) models and results in scientific workflows that incorporate such ML/AI predictions is driven by numerous factors. An uncertainty-aware metric that…

Machine Learning · Computer Science 2023-08-28 Line Pouchard , Kristofer G. Reyes , Francis J. Alexander , Byung-Jun Yoon

Resiliency has garnered attention in the management of critical infrastructure as a metric of system performance, but there are significant roadblocks to its implementation in a realistic decision-making framework. Contrasted to risk and…

Systems and Control · Electrical Eng. & Systems 2026-01-08 Vincent P. Paglioni , Graeme Troxell , Aaron Brown , Steve Conrad , Mazdak Arabi

Screening and surveillance are routinely used in medicine for early detection of disease and close monitoring of progression. Biomarkers are one of the primarily tools used for these tasks, but their successful translation to clinical…

One central goal of design of observational studies is to embed non-experimental data into an approximate randomized controlled trial using statistical matching. Despite empirical researchers' best intention and effort to create…

Methodology · Statistics 2022-06-22 Kan Chen , Siyu Heng , Qi Long , Bo Zhang
‹ Prev 1 2 3 10 Next ›