English
Related papers

Related papers: Modeling data with zero inflation and overdispersi…

200 papers

In this paper, a scale mixture of Normal distributions model is developed for classification and clustering of data having outliers and missing values. The classification method, based on a mixture model, focuses on the introduction of…

Machine Learning · Statistics 2017-11-23 G. Revillon , A. Djafari , C. Enderli

In this survey we present an extensive research of the vast literature about the Generalized Lambda Distribution (GLD) and propose a hurdle, or two-way, model whose associated distribution is the GLD in order to meet the demand for a highly…

Applications · Statistics 2019-01-04 Diego Marcondes , Cláudia Peixoto , Ana Carolina Maia

The statistical modeling of discrete extremes has received less attention than their continuous counterparts in the Extreme Value Theory (EVT) literature. One approach to the transition from continuous to discrete extremes is the modeling…

Methodology · Statistics 2024-06-18 Touqeer Ahmad , Carlo Gaetan , Philippe Naveau

Spatiotemporal data analysis with massive zeros is widely used in many areas such as epidemiology and public health. We use a Bayesian framework to fit zero-inflated negative binomial models and employ a set of latent variables from…

Methodology · Statistics 2024-02-08 Qing He , Hsin-Hsiung Huang

The problem of overdispersion in multivariate count data is a challenging issue. Nowadays, it covers a central role mainly due to the relevance of modern technologies data, such as Next Generation Sequencing and textual data from the web or…

Methodology · Statistics 2025-02-24 Noemi Corsini , Cinzia Viroli

In this paper we introduce the zero-adjusted Birnbaum-Saunders regression model. This new model generalizes at least seven Birnbaum-Saunders regression models. The idea of this modeling is mixing a degenerate distribution at zero with a…

Methodology · Statistics 2020-07-27 Vera Tomazella , Juvêncio S. Nobre , Gustavo H. A. Pereira , Manoel Santos-Neto

Over the last decades, the challenges in applied regression and in predictive modeling have been changing considerably: (1) More flexible model specifications are needed as big(ger) data become available, facilitated by more powerful…

Computation · Statistics 2025-10-07 Nikolaus Umlauf , Nadja Klein , Thorsten Simon , Achim Zeileis

Claim frequency data in insurance records the number of claims on insurance policies during a finite period of time. Given that insurance companies operate with multiple lines of insurance business where the claim frequencies on different…

Applications · Statistics 2022-12-05 Pengcheng Zhang , David Pitt , Xueyuan Wu

We propose a new framework for the modelling of count data exhibiting zero inflation (ZI). The main part of this framework includes a new and more general parameterisation for ZI models which naturally includes both over- and…

Methodology · Statistics 2018-05-03 John Haslett , Andrew Parnell , James Sweeney

In microbiome studies, it is of interest to use a sample from a population of microbes, such as the gut microbiota community, to estimate the population proportion of these taxa. However, due to biases introduced in sampling and…

Methodology · Statistics 2022-10-11 Roulan Jiang , Xiang Zhan , Tianying Wang

This research deals with the estimation and imputation of missing data in longitudinal models with a Poisson response variable inflated with zeros. A methodology is proposed that is based on the use of maximum likelihood, assuming that data…

Methodology · Statistics 2024-09-18 D. S. Martinez-Lobo , O. O. Melo , N. A. Cruz

We use the theory of normal variance-mean mixtures to derive a data augmentation scheme for models that include gamma functions. Our methodology applies to many situations in statistics and machine learning, including Multinomial-Dirichlet…

Methodology · Statistics 2021-06-22 Jingyu He , Nicholas Polson , Jianeng Xu

The incorporation of unlabeled data in regression and classification analysis is an increasing focus of the applied statistics and machine learning literatures, with a number of recent examples demonstrating the potential for unlabeled data…

Methodology · Statistics 2009-09-29 Feng Liang , Sayan Mukherjee , Mike West

Selective inference methods are developed for group lasso estimators for use with a wide class of distributions and loss functions. The method includes the use of exponential family distributions, as well as quasi-likelihood modeling for…

Methodology · Statistics 2024-03-28 Yiling Huang , Sarah Pirenne , Snigdha Panigrahi , Gerda Claeskens

The Poisson-gamma state space (PGSS) models have been utilized in the analysis of non-negative integer-valued time series to sequentially obtain closed form filtering and predictive densities. In this study, we show the underlying mechanics…

Methodology · Statistics 2025-12-18 Kaoru Irie , Tevfik Aktekin

A common type of zero-inflated data has certain true values incorrectly replaced by zeros due to data recording conventions (rare outcomes assumed to be absent) or details of data recording equipment (e.g. artificial zeros in gene…

The current high-dimensional linear factor models fail to account for the different types of variables, while high-dimensional nonlinear factor models often overlook the overdispersion present in mixed-type data. However, overdispersion is…

Methodology · Statistics 2024-08-22 Jinyu Nie , Zhilong Qin , Wei Liu

A key problem in computational sustainability is to understand the distribution of species across landscapes over time. This question gives rise to challenging large-scale prediction problems since (i) hundreds of species have to be…

Machine Learning · Computer Science 2020-11-02 Shufeng Kong , Junwen Bai , Jae Hee Lee , Di Chen , Andrew Allyn , Michelle Stuart , Malin Pinsky , Katherine Mills , Carla P. Gomes

An extensive body of literature exists that specifically addresses the univariate case of zero-inflated count models. In contrast, research pertaining to multivariate models is notably less developed. We proposed two new parsimonious…

Methodology · Statistics 2024-01-17 Claire Geldenhuys , Rene Ehlers , Andriette Bekker

Graphical models are commonly used tools for modeling multivariate random variables. While there exist many convenient multivariate distributions such as Gaussian distribution for continuous data, mixed data with the presence of discrete…

Machine Learning · Statistics 2014-04-30 Jianqing Fan , Han Liu , Yang Ning , Hui Zou