English
Related papers

Related papers: Sequentially valid tests for forecast calibration

200 papers

This paper lays out a principled approach to compare copula forecasts via strictly consistent scores. We first establish the negative result that, in general, copulas fail to be elicitable, implying that copula predictions cannot sensibly…

Methodology · Statistics 2026-02-11 Tobias Fissler , Yannick Hoga

Calibration means that forecasts and average realized frequencies are close. We develop the concept of forecast hedging, which consists of choosing the forecasts so as to guarantee that the expected track record can only improve. This…

Theoretical Economics · Economics 2022-10-14 Dean P. Foster , Sergiu Hart

When machine learning systems meet real world applications, accuracy is only one of several requirements. In this paper, we assay a complementary perspective originating from the increasing availability of pre-trained and regularly…

In situations where forecasters are scored on the quality of their probabilistic predictions, it is standard to use `proper' scoring rules to perform such scoring. These rules are desirable because they give forecasters no incentive to lie…

Methodology · Statistics 2020-08-25 Spencer Greenberg

A probability forecast or probabilistic classifier is reliable or calibrated if the predicted probabilities are matched by ex post observed frequencies, as examined visually in reliability diagrams. The classical binning and counting…

Methodology · Statistics 2021-08-26 Timo Dimitriadis , Tilmann Gneiting , Alexander I. Jordan

We introduce a theoretical framework of elicitability and identifiability of set-valued functionals, such as quantiles, prediction intervals, and systemic risk measures. A functional is elicitable if it is the unique minimiser of an…

Statistics Theory · Mathematics 2022-01-06 Tobias Fissler , Rafael Frongillo , Jana Hlavinová , Birgit Rudloff

We introduce a method for calculating \(p\)-values to test causal hypotheses in qualitative research \emph{a la} process tracing. As in an experiment, our \(p\)-value tells us how often one would make the same or more compelling…

Methodology · Statistics 2025-08-04 Matias Lopez , Jake Bowers

I propose a new procedure to estimate the False Alarm Probability, the measure of significance for peaks of periodograms. The key element of the new procedure is the use of generalized extreme-value distributions, the limiting distribution…

Instrumentation and Methods for Astrophysics · Physics 2012-12-05 M. Süveges

Accurate uncertainty quantification is critical for reliable predictive modeling. Existing methods typically address either aleatoric uncertainty due to measurement noise or epistemic uncertainty resulting from limited data, but not both in…

Machine Learning · Statistics 2026-03-04 Ilia Azizi , Juraj Bodik , Jakob Heiss , Bin Yu

Forecast probabilities often serve as critical inputs for binary decision making. In such settings, calibration$\unicode{x2014}$ensuring forecasted probabilities match empirical frequencies$\unicode{x2014}$is essential. Although the common…

Methodology · Statistics 2025-08-06 Raphael Rossellini , Jake A. Soloff , Rina Foygel Barber , Zhimei Ren , Rebecca Willett

In multivariate time series, the estimation of the covariance matrix of the observation innovations plays an important role in forecasting as it enables the computation of the standardized forecast error vectors as well as it enables the…

Methodology · Statistics 2008-02-04 K. Triantafyllopoulos

Multi-step-ahead forecasts are often updated as new observations become available, since shorter forecast horizons typically improve forecast quality. However, such improvements come at the cost of forecast instability, i.e., variability in…

Machine Learning · Computer Science 2026-05-28 Jente Van Belle , Honglin Wen , Wouter Verbeke , Pierre Pinson

In deploying artificial intelligence (AI) models, selective prediction offers the option to abstain from making a prediction when uncertain about model quality. To fulfill its promise, it is crucial to enforce strict and precise error…

Methodology · Statistics 2026-03-27 Tian Bai , Ying Jin

To mitigate the impacts associated with adverse weather conditions, meteorological services issue weather warnings to the general public. These warnings rely heavily on forecasts issued by underlying prediction systems. When deciding which…

Applications · Statistics 2022-09-13 Sam Allen , Jonas Bhend , Olivia Martius , Johanna Ziegel

Forecast reconciliation is an important research topic. Yet, there is currently neither formal framework nor practical method for the probabilistic reconciliation of count time series. In this paper we propose a definition of coherency and…

Methodology · Statistics 2023-06-28 Giorgio Corani , Dario Azzimonti , Nicolò Rubattu

A storm of favorable or critical publications regarding p-values-based procedures has been observed in both the theoretical and applied literature. We focus on valid definitions of p-values in the scenarios when composite null models are in…

Methodology · Statistics 2020-01-16 Albert Vexler

The usual procedure for estimating the significance of a peak in a power spectrum is to calculate the probability of obtaining that value or a larger value by chance (known as the "p-value"), on the assumption that the time series contains…

High Energy Astrophysical Phenomena · Physics 2009-11-13 P. A. Sturrock , J. D. Scargle

In safety-critical applications data-driven models must not only be accurate but also provide reliable uncertainty estimates. This property, commonly referred to as calibration, is essential for risk-aware decision-making. In regression a…

Machine Learning · Computer Science 2026-04-23 Jelke Wibbeke , Nico Schönfisch , Sebastian Rohjans , Andreas Rauh

In this paper we focus on the problem of assigning uncertainties to single-point predictions generated by a deterministic model that outputs a continuous variable. This problem applies to any state-of-the-art physics or engineering models…

Machine Learning · Statistics 2020-03-12 Enrico Camporeale , Algo Carè

The p-values are often implicitly used as a measure of evidence for the hypotheses of the tests. This practice has been analyzed with different approaches. It is generally accepted for the one-sided hypothesis problem, but it is often…

Statistics Theory · Mathematics 2007-06-13 Guy Morel