English
Related papers

Related papers: A Comparison of Zero-Inflated Models for Modern Bi…

200 papers

Modeling sparse data such as microbiome and transcriptomics (RNA-seq) data is very challenging due to the exceeded number of zeros and skewness of the distribution. Many probabilistic models have been used for modeling sparse data,…

Methodology · Statistics 2021-12-30 Hani Aldirawi , Jie Yang

In epidemiological studies, zero-inflated and hurdle models are commonly used to handle excess zeros in reported infectious disease cases. However, they can not model the persistence (changing from presence to presence) and reemergence…

Applications · Statistics 2025-06-12 Mingchi Xu , Dirk Douwes-Schultz , Alexandra M. Schmidt

Detecting associations between microbial compositions and sample characteristics is one of the most important tasks in microbiome studies. Most of the existing methods apply univariate models to single microbial species separately, with…

The zero-inflated logistic regression model accommodates binary responses with excess zeros, which often arise from a latent mixture of susceptible and insusceptible subpopulations or asymmetric misclassification of the response. The model…

Methodology · Statistics 2026-04-23 Yui Tomo , Shinto Eguchi , Daisuke Yoneoka

Understanding the help and support that is exchanged between family members of different generations is of increasing importance, with research questions in sociology and social policy focusing on both predictors of the levels of help given…

Methodology · Statistics 2025-02-19 Jouni Kuha , Siliang Zhang , Fiona Steele

There are numerous applications which involve modeling multi-dimensional count data, notably in actuarial science and risk management. When such data exhibit an excess of zeros, common count models are no longer suitable. With multivariate…

Methodology · Statistics 2025-09-30 Golshid Aflaki , Juliana Schulz , Jean-François Plante

Single-cell RNA sequencing (scRNA-seq) has revolutionized the study of cellular heterogeneity, enabling detailed molecular profiling at the individual cell level. However, integrating high-dimensional single-cell data into causal mediation…

Methodology · Statistics 2025-10-01 Seungjun Ahn , Li Chen , Maaike van Gerwen , Panos Roussos , Zhigang Li

In many cases, a machine learning model must learn to correctly predict a few data points with particular values of interest in a broader range of data where many target values are zero. Zero-inflated data can be found in diverse scenarios,…

Real-world networks are sparse. As we show in this article, even when a large number of interactions is observed, most node pairs remain disconnected. We demonstrate that classical multi-edge network models, such as the $G(N,p)$,…

Social and Information Networks · Computer Science 2025-01-03 Giona Casiraghi , Georges Andres

Microbiome omics data including 16S rRNA reveal intriguing dynamic associations between the human microbiome and various disease states. Drastic changes in microbiota can be associated with factors like diet, hormonal cycles, diseases, and…

Methodology · Statistics 2024-01-09 Paramahansa Pramanik , Arnab Kumar Maity

Zero-inflated models are frequently used to deal with data having many zeros. A commonly used model for over-dispersed data containing zeros is known as the zero-inflated Poisson model. However, to account for the heterogeneity of counts…

Methodology · Statistics 2025-09-04 Ali Abbas , Sajid Ali , Ismail Shah

Many clinical endpoint measures, such as the number of standard drinks consumed per week or the number of days that patients stayed in the hospital, are count data with excessive zeros. However, the zero-inflated nature of such outcomes is…

Applications · Statistics 2022-07-14 Zhengyang Zhou , Minge Xie , David Huh , Eun-Young Mun

Ecological studies involving counts of abundance, presence-absence or occupancy rates often produce data having a substantial proportion of zeros. Furthermore, these types of processes are typically multivariate and only adequately…

Methodology · Statistics 2011-05-17 Ali Arab , Scott H. Holan , Christopher K. Wikle , Mark L. Wildhaber

We propose a new framework for the modelling of count data exhibiting zero inflation (ZI). The main part of this framework includes a new and more general parameterisation for ZI models which naturally includes both over- and…

Methodology · Statistics 2018-05-03 John Haslett , Andrew Parnell , James Sweeney

Models such as the zero-inflated and zero-altered Poisson and zero-truncated binomial are well-established in modern regression analysis. We propose a super model that jointly and maximally unifies alteration, inflation, truncation and…

Methodology · Statistics 2022-08-30 Thomas W. Yee , Chenchen Ma

Researchers are often interested in predicting outcomes, conducting clustering analysis to detect distinct subgroups of their data, or computing causal treatment effects. Pathological data distributions that exhibit skewness and…

Methodology · Statistics 2020-08-24 Arman Oganisian , Nandita Mitra , Jason Roy

In microbiome studies, it is of interest to use a sample from a population of microbes, such as the gut microbiota community, to estimate the population proportion of these taxa. However, due to biases introduced in sampling and…

Methodology · Statistics 2022-10-11 Roulan Jiang , Xiang Zhan , Tianying Wang

This research deals with the estimation and imputation of missing data in longitudinal models with a Poisson response variable inflated with zeros. A methodology is proposed that is based on the use of maximum likelihood, assuming that data…

Methodology · Statistics 2024-09-18 D. S. Martinez-Lobo , O. O. Melo , N. A. Cruz

Modern RNA sequencing technologies provide gene expression measurements from single cells that promise refined insights on regulatory relationships among genes. Directed graphical models are well-suited to explore such (cause-effect)…

Methodology · Statistics 2020-04-09 Shiqing Yu , Mathias Drton , Ali Shojaie

Count data are ubiquitous in ecology and the Poisson generalized linear model (GLM) is commonly used to model the association between counts and explanatory variables of interest. When fitting this model to the data, one typically proceeds…

Methodology · Statistics 2020-07-14 Harlan Campbell