English
Related papers

Related papers: Simulation Studies For Goodness-of-Fit and Two-Sam…

200 papers

Goodness-of-fit testing is often criticized for its lack of practical relevance: since ``all models are wrong'', the null hypothesis that the data conform to our model is ultimately always rejected as the sample size grows. Despite this,…

Machine Learning · Statistics 2025-10-24 Xing Liu , François-Xavier Briol

Logistic regression is widely used to model the propensity score in the analysis of nonignorable missing data. However, goodness-of-fit testing for this propensity score model has received limited attention in the literature. In this paper,…

Methodology · Statistics 2026-04-24 Manli Cheng , Yangjianchen Xu , Qinglong Tian , Pengfei Li

Exact null distributions of goodness-of-fit test statistics are generally challenging to obtain in tractable forms. Practitioners are therefore usually obliged to rely on asymptotic null distributions or Monte Carlo methods, either in the…

Methodology · Statistics 2023-09-06 Alberto Fernández-de-Marcos , Eduardo García-Portugués

We propose a two-sample test for high-dimensional means that requires neither distributional nor correlational assumptions, besides some weak conditions on the moments and tail properties of the elements in the random vectors. This…

Methodology · Statistics 2019-04-17 Kaijie Xue , Fang Yao

Model checking plays an important role in linear regression as model misspecification seriously affects the validity and efficiency of regression analysis. In practice, model checking is often performed by subjectively evaluating the plot…

Statistics Theory · Mathematics 2019-11-19 Rok Blagus , Jakob Peterlin , Janez Stare

Kernel two-sample tests have been widely used for multivariate data to test equality of distributions. However, existing tests based on mapping distributions into a reproducing kernel Hilbert space mainly target specific alternatives and do…

Methodology · Statistics 2023-11-21 Hoseung Song , Hao Chen

This paper deals with two-sample tests for functional time series data, which have become widely available in conjunction with the advent of modern complex observation systems. Here, particular interest is in evaluating whether two sets of…

Statistics Theory · Mathematics 2019-09-16 Alexander Aue , Holger Dette , Gregory Rice

We introduce a new statistical test based on the observed spacings of ordered data. The statistic is sensitive to detect non-uniformity in random samples, or short-lived features in event time series. Under some conditions, this new test…

Methodology · Statistics 2022-10-27 Philipp Eller , Lolian Shtembari

This paper proposes several tests of restricted specification in nonparametric instrumental regression. Based on series estimators, test statistics are established that allow for tests of the general model against a parametric or…

Econometrics · Economics 2019-09-24 Christoph Breunig

In population genetics and other application fields, models with intractable likelihood are common. Approximate Bayesian Computation (ABC) or more generally Simulation-Based Inference (SBI) methods work by simulating instrumental data sets…

Methodology · Statistics 2025-01-29 Guillaume Le Mailloux , Paul Bastide , Jean-Michel Marin , Arnaud Estoup

Following the line of classification-based two-sample testing, tests based on the Random Forest classifier are proposed. The developed tests are easy to use, require almost no tuning, and are applicable for any distribution on…

Methodology · Statistics 2021-05-07 Simon Hediger , Loris Michel , Jeffrey Näf

This paper introduces a novel goodness-of-fit test technique for parametric conditional distributions. The proposed tests are based on a residual marked empirical process, for which we develop a conditional Principal Component Analysis. The…

Econometrics · Economics 2025-06-18 Cui Rui , Li Yuhao

Due to the broad applications of elliptical models, there is a long line of research on goodness-of-fit tests for empirically validating them. However, the existing literature on this topic is generally confined to low-dimensional settings,…

Statistics Theory · Mathematics 2025-03-04 Siyao Wang , Miles E. Lopes

The Newcomb-Benford probability distribution is becoming very popular in many areas using statistics, notably in fraud detection. In such contexts, it is important to be able to determine if a data set arises from this distribution while…

Statistics Theory · Mathematics 2020-03-03 G. R. Ducharme , S. Kaci , C. Vovor-Dassu

Generalized linear models (GLMs) are used within a vast number of application domains. However, formal goodness of fit (GOF) tests for the overall fit of the model$-$so-called "global" tests$-$seem to be in wide use only for certain classes…

Methodology · Statistics 2021-03-01 Nikola Surjanovic , Richard Lockhart , Thomas M. Loughin

We describe a unified framework within which we can build survival models. The motivation for this work comes from a study on the prediction of relapse among breast cancer patients treated at the Curie Institute in Paris, France. Our focus…

Methodology · Statistics 2014-05-28 Cécile Chauvel , John O'Quigley

It is of great interest to test the equality of the means in two samples of functional data. Past research has predominantly concentrated on low-dimensional functional data, a focus that may not hold up in high-dimensional scenarios. In…

Methodology · Statistics 2025-05-27 Shouxia Wang , Jiguo Cao , Hua Liu , Jinhong You , Jicai Liu

Networks describe the, often complex, relationships between individual actors. In this work, we address the question of how to determine whether a parametric model, such as a stochastic block model or latent space model, fits a dataset well…

Methodology · Statistics 2025-09-10 Shane Lubold , Bolun Liu , Tyler H. McCormick

Goodness-of-fit tests based on the Euclidean distance often outperform chi-square and other classical tests (including the standard exact tests) by at least an order of magnitude when the model being tested for goodness-of-fit is a discrete…

Methodology · Statistics 2024-04-09 William Perkins , Mark Tygert , Rachel Ward

A common disadvantage in existing distribution-free two-sample testing approaches is that the computational complexity could be high. Specifically, if the sample size is $N$, the computational complexity of those two-sample tests is at…

Methodology · Statistics 2017-07-18 Cheng Huang , Xiaoming Huo