English
Related papers

Related papers: A Novel Test of Missing Completely at Random: U-st…

200 papers

Exact tests greatly improve the analysis of contingency tables when marginals are low. For instance, researchers often use Fisher's exact test, which is conditional, or Barnard's test, which is unconditional but needs to deal with a…

Methodology · Statistics 2021-10-28 Miguel Araujo-Voces , Víctor Quesada

Semi-supervised datasets are ubiquitous across diverse domains where obtaining fully labeled data is costly or time-consuming. The prevalence of such datasets has consistently driven the demand for new tools and methods that exploit the…

Statistics Theory · Mathematics 2024-03-12 Ilmun Kim , Larry Wasserman , Sivaraman Balakrishnan , Matey Neykov

We consider identification and estimation with an outcome missing not at random (MNAR). We study an identification strategy based on a so-called shadow variable. A shadow variable is assumed to be correlated with the outcome, but…

Methodology · Statistics 2019-09-10 Wang Miao , Lan Liu , Eric Tchetgen Tchetgen , Zhi Geng

Not all experiments publish their results with a description of the correlations between the data points. This makes it difficult to do hypothesis tests or model fits with that data, since just assuming no correlation can lead to an over-…

Data Analysis, Statistics and Probability · Physics 2021-06-30 Lukas Koch

A very simple interpretation of matrix completion problem is introduced based on statistical models. Combined with the well-known results from missing data analysis, such interpretation indicates that matrix completion is still a valid and…

Machine Learning · Statistics 2016-05-11 Tianxi Li

Two-sample tests for multivariate data and especially for non-Euclidean data are not well explored. This paper presents a novel test statistic based on a similarity graph constructed on the pooled observations from the two samples. It can…

Methodology · Statistics 2024-08-12 Hao Chen , Jerome H. Friedman

We compare two deletion-based methods for dealing with the problem of missing observations in linear regression analysis. One is the complete-case analysis (CC, or listwise deletion) that discards all incomplete observations and only uses…

Methodology · Statistics 2023-05-02 Tianchen Xu , Kun Chen , Gen Li

The large-sample behavior of non-degenerate multivariate $U$-statistics of arbitrary degree is investigated under the assumption that their kernel depends on parameters that can be estimated consistently. Mild regularity conditions are…

Statistics Theory · Mathematics 2025-07-22 Alain Desgagné , Christian Genest , Frédéric Ouimet

Many studies include a goal of determining whether there is treatment effect heterogeneity across different subpopulations. In this paper, we propose a U-statistic-based non-parametric test of the null hypothesis that the treatment effects…

Methodology · Statistics 2020-12-08 Maozhu Dai , Hal S. Stern

This paper tackles the problem of constructing a non-parametric predictor when the latent variables are given with incomplete information. The convenient predictor for this task is the random forest algorithm in conjunction to the so-called…

Statistics Theory · Mathematics 2023-09-01 Irving Gómez-Méndez , Emilien Joly

U-statistics are widely used in fields such as economics, machine learning, and statistics. However, while they enjoy desirable statistical properties, they have an obvious drawback in that the computation becomes impractical as the data…

Statistics Theory · Mathematics 2020-08-12 Xiangshun Kong , Wei Zheng

This paper presents the asymptotic theory for nondegenerate $U$-statistics of high frequency observations of continuous It\^{o} semimartingales. We prove uniform convergence in probability and show a functional stable central limit theorem…

Probability · Mathematics 2014-09-10 Mark Podolskij , Christian Schmidt , Johanna F. Ziegel

A non parametric method based on the empirical likelihood is proposed for detecting the change in the coefficients of high-dimensional linear model where the number of model variables may increase as the sample size increases. This amounts…

Statistics Theory · Mathematics 2015-06-22 Gabriela Ciuperca , Zahraa Salloum

This paper proposes a Kolmogorov-Smirnov type statistic and a Cram\'er-von Mises type statistic to test linearity in semi-functional partially linear regression models. Our test statistics are based on a residual marked empirical process…

Statistics Theory · Mathematics 2022-12-02 Yongzhen Feng , Jie Li , Xiaojun Song

Multiple matrix sampling is a survey methodology technique that randomly chooses a relatively small subset of items to be presented to survey respondents for the purpose of reducing respondent burden. The data produced are missing…

Methodology · Statistics 2017-10-03 Stanislav Kolenikov , Heather Hammer

Missing data occur frequently in empirical studies in health and social sciences, often compromising our ability to make accurate inferences. An outcome is said to be missing not at random (MNAR) if, conditional on the observed variables,…

Methodology · Statistics 2019-01-23 BaoLuo Sun , Lan Liu , Wang Miao , Kathleen Wirth , James Robins , Eric Tchetgen Tchetgen

We introduce fully nonparametric two-sample tests for testing the null hypothesis that the samples come from the same distribution if the values are only indirectly given via current status censoring. The tests are based on the likelihood…

Statistics Theory · Mathematics 2013-07-12 Piet Groeneboom

Missing data is a crucial issue when applying machine learning algorithms to real-world datasets. Starting from the simple assumption that two batches extracted randomly from the same dataset should share the same distribution, we leverage…

Machine Learning · Statistics 2020-07-02 Boris Muzellec , Julie Josse , Claire Boyer , Marco Cuturi

In the missing data literature, the Maximum Likelihood Estimator (MLE) is celebrated for its ignorability property under missing at random (MAR) data. However, its sensitivity to misspecification of the (complete) data model, even under…

Methodology · Statistics 2025-09-23 Badr-Eddine Chérief-Abdellatif , Jeffrey Näf

We offer a natural and extensible measure-theoretic treatment of missingness at random. Within the standard missing data framework, we give a novel characterisation of the observed data as a stopping-set sigma algebra. We demonstrate that…

Methodology · Statistics 2018-01-23 Daniel Farewell , Rhian Daniel , Shaun Seaman