English
Related papers

Related papers: A more interpretable regression model for count da…

200 papers

In this paper, a new mixed Poisson distribution is introduced. This new distribution is obtained by utilizing mixing process, with Poisson distribution as mixed distribution and Transmuted Exponential distribution as mixing distribution.…

Methodology · Statistics 2016-10-05 Deepesh Bhati , Pooja Kumawat , E. Gómez Déniz

Many data sets cannot be accurately described by standard probability distributions due to the excess number of zero values present. For example, zero-inflation is prevalent in microbiome data and single-cell RNA sequencing data, which…

Methodology · Statistics 2024-11-20 Max Beveridge , Zach Goldstein , Hee Cheol Chung

We propose a new methodology to detect zero-inflation and overdispersion based on the comparison of the expected sample extremes among convexly ordered distributions. The method is very flexible and includes tests for the proportion of…

Methodology · Statistics 2008-09-25 A. Baillo , J. Carcamo , J. R. Berrendero

The scan statistic is widely used in spatial cluster detection applications of inhomogeneous Poisson processes. However, real data may present substantial departure from the underlying Poisson process. One of the possible departures has to…

Methodology · Statistics 2013-11-19 André L. F. Cançado , Cibele Q. da-Silva , Michel F. da Silva

We consider three new classes of exponential dispersion models of discrete probability distributions which are defined by specifying their variance functions in their mean value parameterization. In a previous paper (Bar-Lev and Ridder,…

Methodology · Statistics 2020-04-01 Shaul K. Bar-Lev , Ad Ridder

Count data with zero inflation and large outliers are ubiquitous in many scientific applications. However, posterior analysis under a standard statistical model, such as Poisson or negative binomial distribution, is sensitive to such…

Methodology · Statistics 2024-05-09 Yasuyuki Hamura , Kaoru Irie , Shonosuke Sugasawa

This paper proposes a computationally efficient Bayesian factor model for multiple grouped count data. Adopting the link function approach, the proposed model can capture the association within and between the at-risk probabilities and…

Methodology · Statistics 2024-05-13 Genya Kobayashi , Yuta Yamauchi

Models such as the zero-inflated and zero-altered Poisson and zero-truncated binomial are well-established in modern regression analysis. We propose a super model that jointly and maximally unifies alteration, inflation, truncation and…

Methodology · Statistics 2022-08-30 Thomas W. Yee , Chenchen Ma

Ecological studies involving counts of abundance, presence-absence or occupancy rates often produce data having a substantial proportion of zeros. Furthermore, these types of processes are typically multivariate and only adequately…

Methodology · Statistics 2011-05-17 Ali Arab , Scott H. Holan , Christopher K. Wikle , Mark L. Wildhaber

We propose a unified probabilistic framework for sparse count tensors with excess zeros, motivated by single-cell Hi-C data. The observed data are naturally represented as a three-way tensor indexed by genomic loci pairs and cells,…

Methodology · Statistics 2026-04-27 Elena Tuzhilina , Yaoming Zhen

In the analysis of count data often the equidispersion assumption is not suitable, hence the Poisson regression model is inappropriate. As a generalization of the Poisson distribution, the COM-Poisson distribution can deal with under-,…

A frequent challenge encountered with compositional ecological data is how to interpret and model data with a high proportion of zeros and $N$'s. Such data frequently occur in ecological applications where counts of species are collected…

Methodology · Statistics 2025-08-04 James Sweeney , John Haslett , Dipankar Bandyopadhyay , Michael Fop , Andrew C. Parnell

Generalized linear models (GLMs) using a regression procedure to fit relationships between predictor and target variables are widely used in automobile insurance data. Here, in the process of ratemaking and in order to compute the premiums…

Applications · Statistics 2016-06-02 J. M. Pérez-Sánchez , E. Gómez-Déniz

Modeling data with multivariate count responses is a challenging problem due to the discrete nature of the responses. Existing methods for univariate count responses cannot be easily extended to the multivariate case since the dependency…

Methodology · Statistics 2016-08-15 Hao Wu , Xinwei Deng , Naren Ramakrishnan

Several statistical models are given in the form of unnormalized densities, and calculation of the normalization constant is intractable. We propose estimation methods for such unnormalized models with missing data. The key concept is to…

Machine Learning · Statistics 2020-06-11 Masatoshi Uehara , Takeru Matsuda , Jae Kwang Kim

Negative binomial regression is commonly employed to analyze overdispersed count data. With small to moderate sample sizes, the maximum likelihood estimator of the dispersion parameter may be subject to a significant bias, that in turn…

Methodology · Statistics 2020-11-06 Euloge Clovis Kenne Pagui , Alessandra Salvan , Nicola Sartori

The workhorse model for zero-truncated count data (y = 1, 2, ...) is the zero-truncated negative binomial (ZTNB) model. We find it should seldom be used. Instead, we recommend the one-inflated zero-truncated negative binomial (OIZTNB) model…

Econometrics · Economics 2025-03-24 Ryan T. Godwin

The analysis of count data is commonly done using Poisson models. Negative binomial models are a straightforward and readily motivated generalization for the case of overdispersed data, i.e., when the observed variance is greater than…

Methodology · Statistics 2016-01-06 Christian Röver , Stefan Andreas , Tim Friede

We propose a new class of discrete generalized linear models based on the class of Poisson-Tweedie factorial dispersion models with variance of the form $\mu + \phi\mu^p$, where $\mu$ is the mean, $\phi$ and $p$ are the dispersion and…

The problem of overdispersion in multivariate count data is a challenging issue. Nowadays, it covers a central role mainly due to the relevance of modern technologies data, such as Next Generation Sequencing and textual data from the web or…

Methodology · Statistics 2025-02-24 Noemi Corsini , Cinzia Viroli