English
Related papers

Related papers: Accuracy, Repeatability, and Reproducibility of Fi…

200 papers

This article proposes omnibus portmanteau tests for contrasting adequacy of time series models. The test statistics are based on combining the autocorrelation function of the conditional residuals, the autocorrelation function of the…

Methodology · Statistics 2024-02-02 Esam Mahdi

Recent research has generated hope that inference scaling, such as resampling solutions until they pass verifiers like unit tests, could allow weaker models to match stronger ones. Beyond inference, this approach also enables training…

Machine Learning · Computer Science 2026-03-27 Benedikt Stroebl , Sayash Kapoor , Arvind Narayanan

The cost of errors related to machine learning classifiers, namely, false positives and false negatives, are not equal and are application dependent. For example, in cybersecurity applications, the cost of not detecting an attack is very…

Machine Learning · Computer Science 2024-08-01 Manish Marwah , Asad Narayanan , Stephan Jou , Martin Arlitt , Maria Pospelova

The growing awareness of safety concerns in large language models (LLMs) has sparked considerable interest in the evaluation of safety. This study investigates an under-explored issue about the evaluation of LLMs, namely the substantial…

Computation and Language · Computer Science 2024-04-02 Yixu Wang , Yan Teng , Kexin Huang , Chengqi Lyu , Songyang Zhang , Wenwei Zhang , Xingjun Ma , Yu-Gang Jiang , Yu Qiao , Yingchun Wang

Forensic examination of evidence like firearms and toolmarks, traditionally involves a visual and therefore subjective assessment of similarity of two questioned items. Statistical models are used to overcome this subjectivity and allow…

Human-Computer Interaction · Computer Science 2021-11-03 Ganesh Krishnan , Heike Hofmann

Following an extensive simulation study comparing the operating characteristics of three different procedures used for establishing equivalence (the frequentist "TOST", the Bayesian "HDI-ROPE", and the Bayes factor interval null procedure),…

Methodology · Statistics 2022-03-15 Harlan Campbell , Paul Gustafson

The proliferation of automatic faithfulness metrics for summarization has produced a need for benchmarks to evaluate them. While existing benchmarks measure the correlation with human judgements of faithfulness on model-generated summaries,…

Computation and Language · Computer Science 2023-06-06 Liang Ma , Shuyang Cao , Robert L. Logan , Di Lu , Shihao Ran , Ke Zhang , Joel Tetreault , Alejandro Jaimes

Estimating probability of failure in aerospace systems is a critical requirement for flight certification and qualification. Failure probability estimation involves resolving tails of probability distribution, and Monte Carlo sampling…

Numerical Analysis · Mathematics 2022-09-22 S. Ashwin Renganathan , Vishwas Rao , Ionel M. Navon

We consider error correction in quantum key distribution. To avoid that Alice and Bob unwittingly end up with different keys precautions must be taken. Before running the error correction protocol, Bob and Alice normally sacrifice some bits…

Quantum Physics · Physics 2014-10-24 Øystein Marøy , Magne Gudmundsen , Lars Lydersen , Johannes Skaar

A key trait of stochastic optimizers is that multiple runs of the same optimizer in attempting to solve the same problem can produce different results. As a result, their performance is evaluated over several repeats, or runs, on the…

Machine Learning · Computer Science 2026-05-18 Moslem Noori , Elisabetta Valiante , Thomas Van Vaerenbergh , Masoud Mohseni , Ignacio Rozada

We combine in a single framework the two complementary benefits of chi^2-template fits and empirical training sets used e.g. in neural nets: chi^2 is more reliable when its probability density functions (PDFs) are inspected for multiple…

Instrumentation and Methods for Astrophysics · Physics 2015-05-13 Christian Wolf

While in-context learning with large language models (LLMs) has shown impressive performance, we have discovered a unique miscalibration behavior where both correct and incorrect predictions are assigned the same level of confidence. We…

Computation and Language · Computer Science 2024-10-04 Wei Cheng , Tianlu Wang , Yanmin Ji , Fan Yang , Keren Tan , Yiyu Zheng

We present a theoretical framework for the analysis of privacy and security tradeoffs in secure biometric authentication systems. We use this framework to conduct a comparative information-theoretic analysis of two biometric systems that…

Information Theory · Computer Science 2011-12-26 Ye Wang , Shantanu Rane , Stark C. Draper , Prakash Ishwar

This paper develops a Bayesian approach for assessing equivalence and non-inferiority hypotheses in two-arm trials using relative belief ratios. A relative belief ratio is a measure of statistical evidence and can indicate evidence either…

Applications · Statistics 2014-01-20 Saman Muthukumarana , Michael Evans

Gate fidelity -- an average fidelity over all possible input states -- is the workhorse metric for benchmarking quantum gates or circuits, yet fault-tolerant quantum computing ultimately depends on the worst-case behavior, typically…

Quantum Physics · Physics 2026-03-10 Kyoungho Cho , Ilkwon Sohn , Yongsoo Hwang , Jeongho Bang

We analyze different types of simulations that applied researchers can use to assess whether their inference methods reliably control false-positive rates. We show that different assessments involve trade-offs, varying in the types of…

Econometrics · Economics 2025-10-03 Bruno Ferman

Recidivism prediction instruments (RPI's) provide decision makers with an assessment of the likelihood that a criminal defendant will reoffend at a future point in time. While such instruments are gaining increasing popularity across the…

Applications · Statistics 2017-03-02 Alexandra Chouldechova

Dense retrievers and rerankers are central to retrieval-augmented generation (RAG) pipelines, where accurately retrieving factual information is crucial for maintaining system trustworthiness and defending against RAG poisoning. However,…

Information Retrieval · Computer Science 2025-08-29 Haoyu Wu , Qingcheng Zeng , Kaize Ding

The error or variability of machine learning algorithms is often assessed by repeatedly re-fitting a model with different weighted versions of the observed data. The ubiquitous tools of cross-validation (CV) and the bootstrap are examples…

Methodology · Statistics 2020-02-10 Ryan Giordano , Will Stephenson , Runjing Liu , Michael I. Jordan , Tamara Broderick

This study independently reproduces the malware detection methodology presented by Felli cious et al. [7], which employs order-invariant API call frequency analysis using Random Forest classification. We utilized the original public dataset…

Cryptography and Security · Computer Science 2026-01-14 Juhani Merilehto