English
Related papers

Related papers: Informative Features for Model Comparison

200 papers

Generalized linear models are flexible tools for the analysis of diverse datasets, but the classical formulation requires that the parametric component is correctly specified and the data contain no atypical observations. To address these…

Methodology · Statistics 2023-04-21 Ioannis Kalogridis , Gerda Claeskens , Stefan Van Aelst

Real-world data typically contain a large number of features that are often heterogeneous in nature, relevance, and also units of measure. When assessing the similarity between data points, one can build various distance measures using…

Machine Learning · Statistics 2022-05-27 Aldo Glielmo , Claudio Zeni , Bingqing Cheng , Gabor Csanyi , Alessandro Laio

We propose a class of nonparametric two-sample tests with a cost linear in the sample size. Two tests are given, both based on an ensemble of distances between analytic functions representing each of the distributions. The first test uses…

Machine Learning · Statistics 2015-06-16 Kacper Chwialkowski , Aaditya Ramdas , Dino Sejdinovic , Arthur Gretton

Measuring inter-dataset similarity is an important task in machine learning and data mining with various use cases and applications. Existing methods for measuring inter-dataset similarity are computationally expensive, limited, or…

Machine Learning · Computer Science 2025-05-06 Muhammad Rajabinasab , Anton D. Lautrup , Arthur Zimek

Real-world data is often incomplete and contains missing values. To train accurate models over real-world datasets, users need to spend a substantial amount of time and resources imputing and finding proper values for missing data items. In…

Machine Learning · Statistics 2024-03-05 Cheng Zhen , Nischal Aryal , Arash Termehchy , Alireza Aghasi , Amandeep Singh Chabada

We present a unified approach to goodness-of-fit testing in $\mathbb{R}^d$ and on lower-dimensional manifolds embedded in $\mathbb{R}^d$ based on sums of powers of weighted volumes of $k$-th nearest neighbor spheres. We prove asymptotic…

Methodology · Statistics 2016-12-21 Bruno Ebner , Norbert Henze , Joseph E. Yukich

There exist a number of tests for assessing the nonparametric heteroscedastic location-scale assumption. Here we consider a goodness-of-fit test for the more general hypothesis of the validity of this model under a parametric functional…

Statistics Theory · Mathematics 2020-01-01 Marie Hušková , Simos G. Meintanis , Charl Pretorius

Probabilistic representation spaces convey information about a dataset and are shaped by factors such as the training data, network architecture, and loss function. Comparing the information content of such spaces is crucial for…

Machine Learning · Computer Science 2025-02-20 Kieran A. Murphy , Sam Dillavou , Dani S. Bassett

Evaluation of generative models is mostly based on the comparison between the estimated distribution and the ground truth distribution in a certain feature space. To embed samples into informative features, previous works often use…

Machine Learning · Computer Science 2022-12-15 Junghyuk Lee , Jun-Hyuk Kim , Jong-Seok Lee

Model selection and assessment with incomplete data pose challenges in addition to the ones encountered with complete data. There are two main reasons for this. First, many models describe characteristics of the complete data, in spite of…

Methodology · Statistics 2008-08-28 Geert Verbeke , Geert Molenberghs , Caroline Beunckens

Generative models, in particular generative adversarial networks (GANs), have received significant attention recently. A number of GAN variants have been proposed and have been utilized in many applications. Despite large strides in terms…

Computer Vision and Pattern Recognition · Computer Science 2018-10-25 Ali Borji

We derive a new discrepancy statistic for measuring differences between two probability distributions based on combining Stein's identity with the reproducing kernel Hilbert space theory. We apply our result to test how well a probabilistic…

Machine Learning · Statistics 2016-07-04 Qiang Liu , Jason D. Lee , Michael I. Jordan

Given an i.i.d. sample $\{(X_i,Y_i)\}_{i \in \{1 \ldots n\}}$ from the random design regression model $Y = f(X) + \epsilon$ with $(X,Y) \in [0,1] \times [-M,M]$, in this paper we consider the problem of testing the (simple) null hypothesis…

Statistics Theory · Mathematics 2015-02-20 Pierpaolo Brutti

We propose a general and relatively simple method for the construction of goodness-of-fit tests on the sphere and the hypersphere. The method is based on the characterization of probability distributions via their characteristic function,…

Statistics Theory · Mathematics 2023-05-25 Bruno Ebner , Norbert Henze , Simos Meintanis

\textbf{Purpose:} Amplitude analysis is a pivotal tool in hadron spectroscopy, fundamentally involving a series of likelihood fits to multi-dimensional experimental distributions. While robust goodness-of-fit tests exist for low-dimensional…

Data Analysis, Statistics and Probability · Physics 2025-12-02 Huoyi Hou , Beijiang Liu

We study a classification problem with three key challenges: pervasive informative missingness, the integration of partial prior expert knowledge into the learning process, and the need for interpretable decision rules. We propose a…

Machine Learning · Statistics 2026-04-17 Shahar Cohen , David M. Steinberg , Yael Radzyner , Yochai Ben Horin

In this paper, a new goodness-of-fit test for a location-scale family based on progressively Type-II censored order statistics is proposed. Using Monte Carlo simulation studies, the present researchers have observed that the proposed test…

Statistics Theory · Mathematics 2017-04-25 Hamzeh Torabi , Sayyed Mahmoud Mirjalili , Hossein Nadeb

This paper discusses asymptotically distribution free tests for the classical goodness-of-fit hypothesis of an error distribution in nonparametric regression models. These tests are based on the same martingale transform of the residual…

Statistics Theory · Mathematics 2009-09-02 Estate V. Khmaladze , Hira L. Koul

Generative adversarial networks (GANs) are one of the greatest advances in AI in recent years. With their ability to directly learn the probability distribution of data, and then sample synthetic realistic data. Many applications have…

Many flexible families of positive random variables exhibit non-closed forms of the density and distribution functions and this feature is considered unappealing for modelling purposes. However, such families are often characterized by a…

Statistics Theory · Mathematics 2025-06-09 Lucio Barabesi , Antonio Di Noia , Marzia Marcheselli , Caterina Pisani , Luca Pratelli