English
Related papers

Related papers: Modeling Zero-Inflated Correlated Dental Data thro…

200 papers

Detecting associations between microbial compositions and sample characteristics is one of the most important tasks in microbiome studies. Most of the existing methods apply univariate models to single microbial species separately, with…

Learning the joint dependence of discrete variables is a fundamental problem in machine learning, with many applications including prediction, clustering and dimensionality reduction. More recently, the framework of copula modeling has…

Machine Learning · Statistics 2013-11-15 Alfredo Kalaitzis , Ricardo Silva

Count data with high frequencies of zeros are found in many areas, specially in biology. Statistical models to analyze such data started to be developed in the 80s and are still a topic of active research. Such models usually assume a…

Applications · Statistics 2018-10-08 Gustavo Thomas , Luiz R. Nakamura , Rafael A. Moral , Clarice G. B. Demétrio

A key challenge in spatial statistics is the analysis for massive spatially-referenced data sets. Such analyses often proceed from Gaussian process specifications that can produce rich and robust inference, but involve dense covariance…

Methodology · Statistics 2019-07-25 Shinichiro Shirota , Andrew O. Finley , Bruce D. Cook , Sudipto Banerjee

Latent Gaussian copula models provide a powerful means to perform multi-view data integration since these models can seamlessly express dependencies between mixed variable types (binary, continuous, zero-inflated) via latent Gaussian…

Computation · Statistics 2022-04-22 Grace Yoon , Christian L. Müller , Irina Gaynanova

This article proposes a graphical model that handles mixed-type, multi-group data. The motivation for such a model originates from real-world observational data, which often contain groups of samples obtained under heterogeneous conditions…

Methodology · Statistics 2023-01-02 Sjoerd Hermes , Joost van Heerwaarden , Pariya Behrouzi

Mixed data refers to a type of data in which variables can be of multiple types, such as continuous, discrete, or categorical. This data is routinely collected in various fields, including healthcare and social sciences. A common goal in…

Methodology · Statistics 2025-05-22 Mauro Florez , Anna Gottard , Carrie McAdams , Michele Guindani , Marina Vannucci

Multivariate mixed-type outcomes are difficult to model jointly, and additional complexity arises when both marginal effects and dependence structures vary with a covariate such as age or time. Existing approaches often impose restrictive…

Methodology · Statistics 2026-04-15 Yujin Jeong , Seonghyun Jeong

Claim frequency data in insurance records the number of claims on insurance policies during a finite period of time. Given that insurance companies operate with multiple lines of insurance business where the claim frequencies on different…

Applications · Statistics 2022-12-05 Pengcheng Zhang , David Pitt , Xueyuan Wu

In biomedical studies, paired survival data arise naturally when two event times are observed within the same subject. Existing statistical models seldom accommodate both cure fractions and complex dependence structures. In this paper, we…

Methodology · Statistics 2026-04-28 Masaki Hino , Shogo Kato , Takeshi Emura

Microbiome research has immense potential for unlocking insights into human health and disease. A common goal in human microbiome research is identifying subgroups of individuals with similar microbial composition that may be linked to…

Methodology · Statistics 2025-08-21 Suppapat Korsurat , Matthew D. Koslovsky

The National Health and Nutrition Examination Survey (NHANES) is a major program of the National Center for Health Statistics, designed to assess the health and nutritional status of adults and children in the United States. The analysis of…

Applications · Statistics 2019-10-15 Ick Hoon Jin , Fang Liu , Evercita C. Eugenio , Kisung You , Suyu Liu

We introduce a novel approach to estimation problems in settings with missing data. Our proposal -- the Correlation-Assisted Missing data (CAM) estimator -- works by exploiting the relationship between the observations with missing features…

Methodology · Statistics 2020-03-02 Timothy I. Cannings , Yingying Fan

Missing data theory deals with the statistical methods in the occurrence of missing data. Missing data occurs when some values are not stored or observed for variables of interest. However, most of the statistical theory assumes that data…

Recent crash frequency studies incorporate spatiotemporal correlations, but these studies have two key limitations: i) none of these studies accounts for temporal variation in model parameters; and ii) Gibbs sampler suffers from convergence…

Applications · Statistics 2020-08-11 Prasad Buddhavarapu , Prateek Bansal , Jorge A. Prozzi

The Gaussian copula is a powerful tool that has been widely used to model spatial and/or temporal correlated data with arbitrary marginal distributions. However, this kind of model can potentially be too restrictive since it expresses a…

Methodology · Statistics 2023-05-30 Moreno Bevilacqua , Eloy Alvarado , Christian Caamaño-Carrillo

In a traditional Gaussian graphical model, data homogeneity is routinely assumed with no extra variables affecting the conditional independence. In modern genomic datasets, there is an abundance of auxiliary information, which often gets…

Methodology · Statistics 2023-08-16 Yabo Niu , Yang Ni , Debdeep Pati , Bani K. Mallick

A common type of zero-inflated data has certain true values incorrectly replaced by zeros due to data recording conventions (rare outcomes assumed to be absent) or details of data recording equipment (e.g. artificial zeros in gene…

We propose a new highly flexible and tractable Bayesian approach to undertake variable selection in non-Gaussian regression models. It uses a copula decomposition for the joint distribution of observations on the dependent variable. This…

Methodology · Statistics 2020-09-07 Nadja Klein , Michael Stanley Smith

We consider learning continuous probabilistic graphical models in the face of missing data. For non-Gaussian models, learning the parameters and structure of such models depends on our ability to perform efficient inference, and can be…

Machine Learning · Computer Science 2012-03-19 Gal Elidan