English
Related papers

Related papers: Confidence limits: what is the problem? Is there t…

200 papers

Probability measures by themselves, are known to be inappropriate for modeling the dynamics of plain belief and their excessively strong measurability constraints make them unsuitable for some representational tasks, e.g. in the context of…

Artificial Intelligence · Computer Science 2013-02-28 Emil Weydert

With the widespread application of Large Language Models (LLMs) to various domains, concerns regarding the trustworthiness of LLMs in safety-critical scenarios have been raised, due to their unpredictable tendency to hallucinate and…

Computation and Language · Computer Science 2024-11-04 Xin Qiu , Risto Miikkulainen

We propose a general method for constructing hypothesis tests and confidence sets that have finite sample guarantees without regularity conditions. We refer to such procedures as "universal." The method is very simple and is based on a…

Statistics Theory · Mathematics 2022-10-21 Larry Wasserman , Aaditya Ramdas , Sivaraman Balakrishnan

When conducting inference on partially identified parameters, confidence regions may cover the whole identified set with a prescribed probability, to which we will refer as set coverage, or they may cover each of its point with a prescribed…

Econometrics · Economics 2021-06-21 Marc Henry , Alexei Onatski

This paper proposes a belief-based framework for social norms in environments where individuals choose a single action. Relaxing the assumption that the appropriateness standard is common knowledge, the framework allows individuals to be…

Theoretical Economics · Economics 2026-04-30 Senran Lin

Linear combinations of multinomial probabilities, such as those resulting from contingency tables, are of use when evaluating classification system performance. While large sample inference methods for these combinations exist, small sample…

Methodology · Statistics 2021-04-20 Katherine A. Batterton , Christine M. Schubert , Richard L. Warr

In certain applications involving the solution of a Bayesian inverse problem, it may not be possible or desirable to evaluate the full posterior, e.g. due to the high computational cost of doing so. This problem motivates the use of…

Statistics Theory · Mathematics 2024-02-27 Han Cheng Lie , T. J. Sullivan , Aretha Teckentrup

Unbinned likelihood fits are frequent in Physics, and often involve complex functions with several components. We discuss the potential pitfalls of situations where the templates used in the fit are not fixed but depend on the event…

Data Analysis, Statistics and Probability · Physics 2014-11-18 Giovanni Punzi

In data mining, when binary prediction rules are used to predict a binary outcome, many performance measures are used in a vast array of literature for the purposes of evaluation and comparison. Some examples include classification…

Machine Learning · Statistics 2025-07-08 Zheng Yuan , Wenxin Jiang

We study the empirical likelihood approach to construct confidence intervals for the optimal value and the optimality gap of a given solution, henceforth quantify the statistical uncertainty of sample average approximation, for optimization…

Methodology · Statistics 2016-10-25 Henry Lam , Enlu Zhou

We connect the power of Confidence Intervals in different Frequentist methods to their reliability. We show that in the case of a bounded parameter a biased method which near the boundary has large power in testing the parameter against…

High Energy Physics - Experiment · Physics 2007-05-23 C. Giunti , M. Laveder

This paper concerns the construction of confidence intervals in standard seroprevalence surveys. In particular, we discuss methods for constructing confidence intervals for the proportion of individuals in a population infected with a…

Applications · Statistics 2021-10-05 Thomas J. DiCiccio , David M. Ritzwoller , Joseph P. Romano , Azeem M. Shaikh

We propose a frequentist testing procedure that maintains a defined coverage and is optimal in the sense that it gives maximal power to detect deviations from a null hypothesis when the alternative to the null hypothesis is sampled from a…

Applications · Statistics 2020-07-07 Christian Bartels , Johanna Mielke , Ekkehard Glimm

We investigate the calibration of large language models' (LLMs') confidence across diverse tasks. The results of our preregistered study show that the current crop of LLMs are, like people, too sure they are right: confidence exceeds…

Artificial Intelligence · Computer Science 2026-05-26 Noam Michael , Daniel BenShushan , Jacob Bien , Don A. Moore

Classically, confidence intervals are required to have consistent coverage across all values of the parameter. However, this will inevitably break down if the underlying estimation procedure is biased. For this reason, many efforts have…

Methodology · Statistics 2025-08-06 Logan Harris , Patrick Breheny

In this paper, we propose a general method for testing composite hypotheses. Our idea is to use confidence limits to define stopping and decision rules. The requirements of operating characteristic function can be satisfied by adjusting the…

Statistics Theory · Mathematics 2012-02-10 Xinjia Chen

In Bayesian statistics, the choice of prior distribution is often debatable, especially if prior knowledge is limited or data are scarce. In imprecise probability, sets of priors are used to accurately model and reflect prior knowledge.…

Methodology · Statistics 2016-10-25 Gero Walter , Frank P. A. Coolen

Most NLP datasets are not annotated with protected attributes such as gender, making it difficult to measure classification bias using standard measures of fairness (e.g., equal opportunity). However, manually annotating a large dataset…

Computation and Language · Computer Science 2020-04-28 Kawin Ethayarajh

Confidence intervals are assessed according to two criteria, namely expected length and coverage probability. In an attempt to apply the decision-theoretic method to finding a good confidence interval, a loss function that is a linear…

Statistics Theory · Mathematics 2017-10-18 Paul Kabaila

Bayes [Philos. Trans. R. Soc. Lond. 53 (1763) 370--418; 54 296--325] introduced the observed likelihood function to statistical inference and provided a weight function to calibrate the parameter; he also introduced a confidence…

Methodology · Statistics 2011-12-26 D. A. S. Fraser