English
Related papers

Related papers: GRASP: A Goodness-of-Fit Test for Classification L…

200 papers

Fitting mixture distributions is needed in applications where data belongs to inhomogeneous populations comprising homogeneous sub-populations. The mixing proportions of the sub populations are in general unknown and need to be estimated as…

Methodology · Statistics 2019-12-10 Richard A. Lockhart , Chandanie W. Navaratna

We develop a general statistical framework for the analysis and inference of large tree-structured data, with a focus on developing asymptotic goodness-of-fit tests. We first propose a consistent statistical model for binary trees, from…

We investigate the problem of classification in the presence of unknown class-conditional label noise in which the labels observed by the learner have been corrupted with some unknown class dependent probability. In order to obtain finite…

Machine Learning · Statistics 2019-06-11 Henry W J Reeve , Ata Kaban

Labeling a classification dataset implies to define classes and associated coarse labels, that may approximate a smoother and more complicated ground truth. For example, natural images may contain multiple objects, only one of which is…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Raphael Baena , Lucas Drumetz , Vincent Gripon

This paper presents and examines computationally convenient goodness-of-fit tests for the family of generalized Poisson distributions, which encompasses notable distributions such as the Compound Poisson and the Katz distributions. The…

Methodology · Statistics 2024-11-21 A. Batsidis , B. Milošević , M. D. Jiménez-Gamero

We consider marked empirical processes indexed by a randomly projected functional covariate to construct goodness-of-fit tests for the functional linear model with scalar response. The test statistics are built from continuous functionals…

Classification is a fundamental problem in machine learning and data mining. During the past decades, numerous classification methods have been presented based on different principles. However, most existing classifiers cast the…

Machine Learning · Computer Science 2019-04-23 Zengyou He , Chaohua Sheng , Yan Liu , Quan Zou

Selective classifiers improve model reliability by abstaining on inputs the model deems uncertain. However, few practical approaches achieve the gold-standard performance of a perfect-ordering oracle that accepts examples exactly in order…

Machine Learning · Computer Science 2025-10-27 Stephan Rabanser , Nicolas Papernot

We present a unified approach to goodness-of-fit testing in $\mathbb{R}^d$ and on lower-dimensional manifolds embedded in $\mathbb{R}^d$ based on sums of powers of weighted volumes of $k$-th nearest neighbor spheres. We prove asymptotic…

Methodology · Statistics 2016-12-21 Bruno Ebner , Norbert Henze , Joseph E. Yukich

We propose a goodness-of-fit test for the distribution of errors from a multivariate indirect regression model. The test statistic is based on the Khmaladze transformation of the empirical process of standardized residuals. This…

Methodology · Statistics 2018-12-07 Justin Chown , Nicolai Bissantz , Holger Dette

The reproducing kernel Hilbert space (RKHS) embedding of distributions offers a general and flexible framework for testing problems in arbitrary domains and has attracted considerable amount of attention in recent years. To gain insights…

Machine Learning · Statistics 2017-09-26 Krishnakumar Balasubramanian , Tong Li , Ming Yuan

This paper develops goodness of fit statistics that can be used to formally assess Markov random field models for spatial data, when the model distributions are discrete or continuous and potentially parametric. Test statistics are formed…

Statistics Theory · Mathematics 2012-05-29 Mark S. Kaiser , Soumendra N. Lahiri , Daniel J. Nordman

We present a novel subset scan method to detect if a probabilistic binary classifier has statistically significant bias -- over or under predicting the risk -- for some subgroup, and identify the characteristics of this subgroup. This form…

Machine Learning · Statistics 2017-07-05 Zhe Zhang , Daniel B. Neill

Conformal prediction constructs a set of labels instead of a single point prediction, while providing a probabilistic coverage guarantee. Beyond the coverage guarantee, adaptiveness to example difficulty is an important property. It means…

Machine Learning · Computer Science 2025-11-18 Sooyong Jang , Insup Lee

The coefficient of determination, known as $R^2$, is commonly used as a goodness-of-fit criterion for fitting linear models. $R^2$ is somewhat controversial when fitting nonlinear models, although it may be generalised on a case-by-case…

Methodology · Statistics 2021-12-23 Mark Levene , Aleksejus Kononovicius

We present the results of a large number of simulation studies regarding the power of various goodness-of-fit as well as nonparametric two-sample tests for univariate data. This includes both continuous and discrete data. In general no…

Methodology · Statistics 2024-11-13 Wolfgang Rolke

This paper proposes a novel two-step strategy for testing the goodness-of-fit of parametric regression models in ultra-high dimensional sparse settings, where the predictor dimension far exceeds the sample size. This regime usually renders…

Methodology · Statistics 2025-12-30 Falong Tan , Jie Liu , Heng Peng , Lixing Zhu

Model checking plays an important role in linear regression as model misspecification seriously affects the validity and efficiency of regression analysis. In practice, model checking is often performed by subjectively evaluating the plot…

Statistics Theory · Mathematics 2019-11-19 Rok Blagus , Jakob Peterlin , Janez Stare

We consider a new group testing model wherein each item is a binary random variable defined by an a priori probability of being defective. We assume that each probability is small and that items are independent, but not necessarily…

Information Theory · Computer Science 2018-07-24 Tongxin Li , Chun Lam Chan , Wenhao Huang , Tarik Kaced , Sidharth Jaggi

We introduce a new goodness-of-fit test for count data on $\mathbb{N}$ for the Zeta distribution with unknown parameter. The test is built on a Stein-type characterization that uses, as Stein operator, the infinitesimal generator of a…

Statistics Theory · Mathematics 2026-01-01 Bruno Ebner , Daniel Hlubinka