English
Related papers

Related papers: A safe Hosmer-Lemeshow test

200 papers

There is a useful counterpart of conformal prediction for e-values, called conformal e-prediction. Conformal prediction can serve as basis for testing the assumption of exchangeability, leading to conformal testing. Similarly, conformal…

Statistics Theory · Mathematics 2024-11-05 Vladimir Vovk , Ilia Nouretdinov , Alex Gammerman

Model calibration aims to align confidence with prediction correctness. The Cross-Entropy (CE) loss is widely used for calibrator training, which enforces the model to increase confidence on the ground truth class. However, we find the CE…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Yuchi Liu , Lei Wang , Yuli Zou , James Zou , Liang Zheng

Quantum Hoare logic (QHL) is a formal verification tool specifically designed to ensure the correctness of quantum programs. There has been an ongoing challenge to achieve a relatively complete satisfaction-based QHL with while-loop since…

Logic in Computer Science · Computer Science 2024-05-06 Xin Sun , Xingchi Su , Xiaoning Bian , Huiwen Wu

Post-hoc multi-class calibration is a common approach for providing high-quality confidence estimates of deep neural network predictions. Recent work has shown that widely used scaling methods underestimate their calibration error, while…

Machine Learning · Computer Science 2022-11-28 Kanil Patel , William Beluch , Bin Yang , Michael Pfeiffer , Dan Zhang

We present a statistical test that can be used to verify supervisory requirements concerning overlapping time windows for the long-term calibration in rating systems. In a first step, we show that the long-run default rate is approximately…

Risk Management · Quantitative Finance 2023-12-25 Patrick Kurth , Max Nendel , Jan Streicher

Universal hypothesis testing refers to the problem of deciding whether samples come from a nominal distribution or an unknown distribution that is different from the nominal distribution. Hoeffding's test, whose test statistic is equivalent…

Information Theory · Computer Science 2017-11-15 Pengfei Yang , Biao Chen

We study the interpretability of conditional probability estimates for binary classification under the agnostic setting or scenario. Under the agnostic setting, conditional probability estimates do not necessarily reflect the true…

Machine Learning · Computer Science 2017-03-01 Yihan Gao , Aditya Parameswaran , Jian Peng

Many studies have observed that modern neural networks achieve high accuracy while producing poorly calibrated probabilities, making calibration a critical practical issue. In this work, we propose probability bounding (PB), a novel…

Machine Learning · Statistics 2026-02-24 Kyohei Atarashi , Satoshi Oyama , Hiromi Arai , Hisashi Kashima

We introduce $\textit{Backward Conformal Prediction}$, a method that guarantees conformal coverage while providing flexible control over the size of prediction sets. Unlike standard conformal prediction, which fixes the coverage level and…

Machine Learning · Statistics 2026-02-13 Etienne Gauthier , Francis Bach , Michael I. Jordan

This paper introduces e-fold cross-validation, an energy-efficient alternative to k-fold cross-validation. It dynamically adjusts the number of folds based on a stopping criterion. The criterion checks after each fold whether the standard…

Machine Learning · Computer Science 2024-10-29 Christopher Mahlich , Tobias Vente , Joeran Beel

Pairwise comparisons are a well-known method for modelling of the subjective preferences of a decision maker. A popular implementation of the method is based on solving an eigenvalue problem for M - the matrix of pairwise comparisons. This…

Discrete Mathematics · Computer Science 2015-09-25 Konrad Kułakowski

A new test is proposed for the weak white noise null hypothesis. The test is based on a new automatic choice of the order for a Box-Pierce or Hong test statistic. The test uses Lobato (2001) or Kuan and Lee (2006) HAC critical values. The…

Statistics Theory · Mathematics 2019-08-16 Alain Guay , Emmanuel Guerre , Stepana Lazarova

In this paper we have updated the hypothesis testing framework by drawing upon modern computational power and classification models from machine learning. We show that a simple classification algorithm such as a boosted decision stump can…

Econometrics · Economics 2021-03-03 Gary Cornwall , Jeff Chen , Beau Sauley

We consider the problem of multiple hypothesis testing with generic side information: for each hypothesis $H_i$ we observe both a p-value $p_i$ and some predictor $x_i$ encoding contextual information about the hypothesis. For large-scale…

Methodology · Statistics 2018-07-26 Lihua Lei , William Fithian

Distribution-free predictive inference beyond the construction of prediction sets has gained a lot of interest in recent applications. One such application is the selection task, where the objective is to design a reliable selection rule to…

Methodology · Statistics 2025-01-07 Yonghoon Lee , Zhimei Ren

As probabilistic models continue to permeate various facets of our society and contribute to scientific advancements, it becomes a necessity to go beyond traditional metrics such as predictive accuracy and error rates and assess their…

Machine Learning · Statistics 2025-05-05 Ritwik Vashistha , Arya Farahi

In Bayesian statistics, the marginal likelihood, also known as the evidence, is used to evaluate model fit as it quantifies the joint probability of the data under the prior. In contrast, non-Bayesian models are typically compared using…

Methodology · Statistics 2019-09-24 Edwin Fong , Chris Holmes

It is often of interest to test a global null hypothesis using multiple, possibly dependent $p$-values by combining their strengths while controlling the type-I error. Recently, several heavy-tailed combination tests, such as the harmonic…

Statistics Theory · Mathematics 2026-03-25 Parijat Chakraborty , F. Richard Guo , Kerby Shedden , Stilian Stoev

For users to trust model predictions, they need to understand model outputs, particularly their confidence - calibration aims to adjust (calibrate) models' confidence to match expected accuracy. We argue that the traditional calibration…

Computation and Language · Computer Science 2022-10-25 Chenglei Si , Chen Zhao , Sewon Min , Jordan Boyd-Graber

Any decision making process that relies on a probabilistic forecast of future events necessarily requires a calibrated forecast. This paper proposes new methods for empirically assessing forecast calibration in a multivariate setting where…

Methodology · Statistics 2014-07-02 Thordis L. Thorarinsdottir , Michael Scheuerer , Christopher Heinz