Related papers: Confidence limits: what is the problem? Is there t…
Motivated by applications to goodness of fit testing, the empirical likelihood approach is generalized to allow for the number of constraints to grow with the sample size and for the constraints to use estimated criteria functions. The…
We consider the problem of constructing confidence intervals for the median of a response $Y \in \mathbb{R}$ conditional on features $X \in \mathbb{R}^d$ in a situation where we are not willing to make any assumption whatsoever on the…
We suggest how to construct joint confidence distributions for several parameters and apply these ideas to an autoregressive process of general order. The implied non informative prior for the parameters, i.e. the ratio between the…
This paper presents an argument for why we are not measuring trust sufficiently in explainability, interpretability, and transparency research. Most studies ask participants to complete a trust scale to rate their trust of a model that has…
A common assumption in belief revision is that the reliability of the information sources is either given, derived from temporal information, or the same for all. This article does not describe a new semantics for integration but the…
We characterize a notion of confidence that arises in learning or updating beliefs: the amount of trust one has in incoming information and its impact on the belief state. This learner's confidence can be used alongside (and is easily…
There are things we know, things we know we don't know, and then there are things we don't know we don't know. In this paper we address the latter two issues in a Bayesian framework, introducing the notion of doubt to quantify the degree of…
Prediction credibility measures, in the form of confidence intervals or probability distributions, are fundamental in statistics and machine learning to characterize model robustness, detect out-of-distribution samples (outliers), and…
Bounded proofs are convenient to use due to the high degree of automation that exhaustive checking affords. However, they fall short of providing the robust assurances offered by unbounded proofs. We sketch how completeness thresholds serve…
Constructing confidence intervals that are simultaneously valid across a class of estimates is central to tasks such as multiple mean estimation, generalization guarantees, and adaptive experimental design. We frame this as an ``error…
Rating systems are ubiquitous, with applications ranging from product recommendation to teaching evaluations. Confidence intervals for functionals of rating data such as empirical means or quantiles are critical to decision-making in…
Bounded confidence opinion dynamics model the propagation of information in social networks. However in the existing literature, opinions are only viewed as abstract quantities without semantics rather than as part of a decision-making…
This paper is meant as a contribution to the often debated subject of how to combine data which appear to be in mutual disagreement. As a practical example, the epsilon-prime/epsilon determinations have been considered.
There has not been an established mathematical measure of evidence. Some Bayesians have argued that probability can be an objectively correct measure of ``rational degrees of belief,'' which we do not distinguish from evidence. However,…
Is it possible for a large sequence of measurements or observations, which support a hypothesis, to counterintuitively decrease our confidence? Can unanimous support be too good to be true? The assumption of independence is often made in…
Null hypothesis significance testing remains popular despite decades of concern about misuse and misinterpretation. We believe that much of the problem is due to language: significance testing has little to do with other meanings of the…
We propose using a Bayes procedure with uniform improper prior to determine credible belts for the mean of a Poisson distribution in the presence of background and for the continuous problem of measuring a non-negative quantity $\theta$…
Fisher's fiducial probability has recently received renewed attention under the name confidence. In this paper, we reformulate it within an extended-likelihood framework, a representation that helps to resolve many long-standing…
There is a growing interest in societal concerns in machine learning systems, especially in fairness. Multicalibration gives a comprehensive methodology to address group fairness. In this work, we address the multicalibration error and…
In statistical inference, confidence set procedures are typically evaluated based on their validity and width properties. Even when procedures achieve rate-optimal widths, confidence sets can still be excessively wide in practice due to…