English
Related papers

Related papers: Multiple Imputation Using Gaussian Copulas

200 papers

In this work we present a rigorous application of the Expectation Maximization algorithm to determine the marginal distributions and the dependence structure in a Gaussian copula model with missing data. We further show how to circumvent a…

Machine Learning · Statistics 2022-01-17 Maximilian Kertel , Markus Pauly

We propose a method for post-processing an ensemble of multivariate forecasts in order to obtain a joint predictive distribution of weather. Our method utilizes existing univariate post-processing techniques, in this case ensemble Bayesian…

Applications · Statistics 2015-10-28 Annette Möller , Alex Lenkoski , Thordis L. Thorarinsdottir

This article proposes a graphical model that handles mixed-type, multi-group data. The motivation for such a model originates from real-world observational data, which often contain groups of samples obtained under heterogeneous conditions…

Methodology · Statistics 2023-01-02 Sjoerd Hermes , Joost van Heerwaarden , Pariya Behrouzi

Missing data are ubiquitous in real world applications and, if not adequately handled, may lead to the loss of information and biased findings in downstream analysis. Particularly, high-dimensional incomplete data with a moderate sample…

Machine Learning · Computer Science 2022-12-23 Zongyu Dai , Zhiqi Bu , Qi Long

Multiple imputation is a highly recommended technique to deal with missing data, but the application to longitudinal datasets can be done in multiple ways. When a new wave of longitudinal data arrives, we can treat the combined data of…

Methodology · Statistics 2026-05-18 X. M. Kavelaars , S. van Buuren , J. R. van Ginkel

The imputation of the Multivariate time series (MTS) is particularly challenging since the MTS typically contains irregular patterns of missing values due to various factors such as instrument failures, interference from irrelevant data,…

Machine Learning · Computer Science 2025-04-04 Ye Su , Hezhe Qiao , Di Wu , Yuwen Chen , Lin Chen

This paper proposes a flexible Bayesian approach to multiple imputation using conditional Gaussian mixtures. We introduce novel shrinkage priors for covariate-dependent mixing proportions in the mixture models to automatically select the…

Methodology · Statistics 2022-08-17 Shonosuke Sugasawa , Jae Kwang Kim , Kosuke Morikawa

Copula-based methods provide a flexible approach to build missing data imputation models of multivariate data of mixed types. However, the choice of copula function is an open question. We consider a Bayesian nonparametric approach by using…

Methodology · Statistics 2019-10-15 Jiali Wang , Anton Westveld , Bronwyn Loong , Alan Welsh

International comparisons of hierarchical time series data sets based on survey data, such as annual country-level estimates of school enrollment rates, can suffer from large amounts of missing data due to differing coverage of surveys…

Methodology · Statistics 2025-03-31 Daphne H. Liu , Adrian E. Raftery

We propose a novel distributional regression model for a multivariate response vector based on a copula process over the covariate space. It uses the implicit copula of a Gaussian multivariate regression, which we call a ``regression…

Methodology · Statistics 2024-03-06 Nadja Klein , Michael Stanley Smith , David Nott , Ryan Chisholm

Zero-inflated continuous data ubiquitously appear in many fields, in which lots of exactly zero-valued data are observed while others distribute continuously. Due to the mixed structure of discreteness and continuity in its distribution,…

Methodology · Statistics 2024-10-28 Keita Hamamoto

Estimating copulas with discrete marginal distributions is challenging, especially in high dimensions, because computing the likelihood contribution of each observation requires evaluating $2^{J}$ terms, with $J$ the number of discrete…

Methodology · Statistics 2018-11-12 D. Gunawan , M. -N. Tran , K. Suzuki , J. Dick , R. Kohn

We propose a new copula model for replicated multivariate spatial data. Unlike classical models that assume multivariate normality of the data, the proposed copula is based on the assumption that some factors exist that affect the joint…

Applications · Statistics 2018-10-12 Pavel Krupskii , Marc G. Genton

We discuss the connection between information and copula theories by showing that a copula can be employed to decompose the information content of a multivariate distribution into marginal and dependence components, with the latter…

Statistical Finance · Quantitative Finance 2011-10-26 Rafael S. Calsaverini , Renato Vicente

We propose a multiple imputation method based on principal component analysis (PCA) to deal with incomplete continuous data. To reflect the uncertainty of the parameters from one imputation to the next, we use a Bayesian treatment of the…

Methodology · Statistics 2015-08-20 Vincent Audigier , François Husson , Julie Josse

The problem of missing values in multivariable time series is a key challenge in many applications such as clinical data mining. Although many imputation methods show their effectiveness in many applications, few of them are designed to…

Machine Learning · Computer Science 2020-03-04 Ye Xue , Diego Klabjan , Yuan Luo

Multivariate datasets are common in various real-world applications. Recently, copulas have received significant attention for modeling dependencies among random variables. A copula-based information measure is required to quantify the…

Methodology · Statistics 2024-08-06 Mohd. Arshad , Swaroop Georgy Zachariah , Ashok Kumar Pathak

Gaussian copulas are widely used to estimate multivariate distributions and relationships. We present algorithms for estimating Gaussian copula correlations that ensure differential privacy. We first convert data values into sets of two-way…

Methodology · Statistics 2026-01-08 Shuo Wang , Joseph Feldman , Jerome P. Reiter

We consider the situation of estimating Cox regression in which some covariates are subject to missing, and there exists additional information (including observed event time, censoring indicator and fully observed covariates) which may be…

Methodology · Statistics 2017-10-16 Chiu-Hsieh Hsu , Mandi Yu

Non-random sample selection is a commonplace amongst many empirical studies and it appears when an output variable of interest is available only for a restricted non-random sub-sample of data. We introduce an extension of the generalized…

Statistics Theory · Mathematics 2015-08-18 M. Wojtyś , G. Marra