English
Related papers

Related papers: On Testability and Goodness of Fit Tests in Missin…

200 papers

Many real-world classification problems are significantly class-imbalanced to detriment of the class of interest. The standard set of proper evaluation metrics is well-known but the usual assumption is that the test dataset imbalance equals…

Machine Learning · Computer Science 2020-04-16 Jan Brabec , Tomáš Komárek , Vojtěch Franc , Lukáš Machlica

Model selection and assessment with incomplete data pose challenges in addition to the ones encountered with complete data. There are two main reasons for this. First, many models describe characteristics of the complete data, in spite of…

Methodology · Statistics 2008-08-28 Geert Verbeke , Geert Molenberghs , Caroline Beunckens

Though many deep learning-based models have made great progress in vulnerability detection, we have no good understanding of these models, which limits the further advancement of model capability, understanding of the mechanism of model…

Software Engineering · Computer Science 2024-08-15 Chao Ni , Liyu Shen , Xiaodan Xu , Xin Yin , Shaohua Wang

Testing procedures for assessing a parametric regression model with circular response and $\mathbb{R}^d$-valued covariate are proposed and analyzed in this work both for independent and for spatially correlated data. The test statistics are…

Methodology · Statistics 2020-09-01 Andrea Meilán-Vila , Mario Francisco-Fernández , Rosa M. Crujeiras

Empirical studies suggest that machine learning models often rely on features, such as the background, that may be spuriously correlated with the label only during training time, resulting in poor accuracy during test-time. In this work, we…

Machine Learning · Computer Science 2024-09-10 Vaishnavh Nagarajan , Anders Andreassen , Behnam Neyshabur

When deployed in the real world, machine learning models inevitably encounter changes in the data distribution, and certain -- but not all -- distribution shifts could result in significant performance degradation. In practice, it may make…

Machine Learning · Statistics 2022-05-06 Aleksandr Podkopaev , Aaditya Ramdas

The chain graph model admits both undirected and directed edges in one graph, where symmetric conditional dependencies are encoded via undirected edges and asymmetric causal relations are encoded via directed edges. Though frequently…

Methodology · Statistics 2024-01-29 Ruixuan Zhao , Haoran Zhang , Junhui Wang

Deep network models perform excellently on In-Distribution (ID) data, but can significantly fail on Out-Of-Distribution (OOD) data. While developing methods focus on improving OOD generalization, few attention has been paid to evaluating…

Machine Learning · Computer Science 2021-11-22 Rui Hu , Jitao Sang , Jinqiang Wang , Rui Hu , Chaoquan Jiang

Mixed-effects logistic regression is widely used for binary outcomes in hierarchical data, yet formal goodness-of-fit tests remain limited to random-intercept models and do not address sparse cluster settings. We extend a grouping-based…

Methodology · Statistics 2026-04-22 Ariel Linden

Generalized Linear Models (GLMs) are an increasingly popular framework for modeling neural spike trains. They have been linked to the theory of stochastic point processes and researchers have used this relation to assess goodness-of-fit…

Neurons and Cognition · Quantitative Biology 2010-11-19 Felipe Gerhard , Wulfram Gerstner

In this paper we present the methodology for detecting outliers and testing the goodness-of-fit of random sets using topological data analysis. We construct the filtration from level sets of the signed distance function and consider various…

Methodology · Statistics 2024-04-09 Vesna Gotovac Đogaš , Marcela Mandarić

Models for analyzing multivariate data sets with missing values require strong, often unassessable, assumptions. The most common of these is that the mechanism that created the missing data is ignorable - a twofold assumption dependent on…

Applications · Statistics 2020-02-17 Iavor Bojinov , Natesh Pillai , Donald Rubin

Data analyses typically rely upon assumptions about missingness mechanisms that lead to observed versus missing data. When the data are missing not at random, direct assumptions about the missingness mechanism, and indirect assumptions…

Methodology · Statistics 2016-03-22 Alexander M Franks , Edoardo M Airoldi , Donald B Rubin

Data depth provides a centre-outward ordering for multivariate data. Recently, some univariate GoF tests based on data depth have been studied by Li (2018). This paper discusses some univariate goodness of fit tests based on centre-outward…

Methodology · Statistics 2024-05-14 Rahul Singh

Linear mixed effects models (LMMs) are a popular and powerful tool for analyzing clustered or repeated observations for numeric outcomes. LMMs consist of a fixed and a random component, specified in the model through their respective design…

Statistics Theory · Mathematics 2019-12-10 Rok Blagus , Jakob Peterlin , Nataša Kejžar

Model misspecification can create significant challenges for the implementation of probabilistic models, and this has led to development of a range of robust methods which directly account for this issue. However, whether these more…

Machine Learning · Statistics 2025-04-22 Oscar Key , Arthur Gretton , François-Xavier Briol , Tamara Fernandez

We describe a unified framework within which we can build survival models. The motivation for this work comes from a study on the prediction of relapse among breast cancer patients treated at the Curie Institute in Paris, France. Our focus…

Methodology · Statistics 2014-05-28 Cécile Chauvel , John O'Quigley

Datasets with missing values are very common on industry applications, and they can have a negative impact on machine learning models. Recent studies introduced solutions to the problem of imputing missing values based on deep generative…

Machine Learning · Computer Science 2019-02-28 Ramiro D. Camino , Christian A. Hammerschmidt , Radu State

Distribution testing can be described as follows: $q$ samples are being drawn from some unknown distribution $P$ over a known domain $[n]$. After the sampling process, a decision must be made about whether $P$ holds some property, or is far…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-12-04 Uri Meir

We consider the goodness of fit testing problem for stochastic differential equation with small diffiusion coefficient. The basic hypothesis is always simple and it is described by the known trend coefficient. We propose several tests of…

Statistics Theory · Mathematics 2009-03-27 Yury A. Kutoyants