English
Related papers

Related papers: Zero-inflated Poisson Factor Model with Applicatio…

200 papers

A common type of zero-inflated data has certain true values incorrectly replaced by zeros due to data recording conventions (rare outcomes assumed to be absent) or details of data recording equipment (e.g. artificial zeros in gene…

Suppose x is any exactly k-sparse vector in R^n. We present a class of sparse matrices A, and a corresponding algorithm that we call SHO-FA (for Short and Fast) that, with high probability over A, can reconstruct x from Ax. The SHO-FA…

Information Theory · Computer Science 2012-11-16 Mayank Bakshi , Sidharth Jaggi , Sheng Cai , Minghua Chen

This work explores non-negative low-rank matrix factorization based on regularized Poisson models (PF or "Poisson factorization" for short) for recommender systems with implicit-feedback data. The properties of Poisson likelihood allow a…

Machine Learning · Computer Science 2022-02-25 David Cortes

The paper proposes a latent variable model for binary data coming from an unobserved heterogeneous population. The heterogeneity is taken into account by replacing the traditional assumption of Gaussian distributed factors by a finite…

Methodology · Statistics 2010-10-13 Silvia Cagnone , Cinzia Viroli

Many data-driven approaches exist to extract neural representations of functional magnetic resonance imaging (fMRI) data, but most of them lack a proper probabilistic formulation. We propose a group level scalable probabilistic sparse…

We propose a regression model for count data when the classical generalized linear model approach is too rigid due to a high outcome of zero counts and a nonlinear influence of continuous covariates. Zero-Inflation is applied to take into…

Methodology · Statistics 2013-04-12 T. Opitz , P. Tramini , N. Molinari

Spatial two-component mixture models offer a robust framework for analyzing spatially correlated data with zero inflation. To circumvent potential biases introduced by assuming a specific distribution for the response variables, we employ a…

Methodology · Statistics 2025-09-17 Chung-Wei Shen , Bu-Ren Hsu , Chia-Ming Hsu , Chun-Shu Chen

Recent advances in engineering technologies have enabled the collection of a large number of longitudinal features. This wealth of information presents unique opportunities for researchers to investigate the complex nature of diseases and…

Methodology · Statistics 2023-11-27 Zihang Lu , Noirrit Kiran Chandra

By creating networks of biochemical pathways, communities of micro-organisms are able to modulate the properties of their environment and even the metabolic processes within their hosts. Next-generation high-throughput sequencing has led to…

Applications · Statistics 2023-03-28 Molly G. Hayes , Morgan G. I. Langille , Hong Gu

We introduce a new class of Poisson-exponential-Tweedie (PET) mixture in the framework of generalized linear models for ultra-overdispersed count data. The mean-variance relationship is of the form $m+m^{2}+\phi m^{p}$, where $\phi$ and $p$…

Methodology · Statistics 2019-08-26 Rahma Abid , Celestin C. Kokonendji , Afif Masmoudi

The human microbiome plays an important role in human health and disease status. Next generating sequencing technologies allow for quantifying the composition of the human microbiome. Clustering these microbiome data can provide valuable…

Methodology · Statistics 2021-01-07 Wangshu Tu , Sanjeena Subedi

Revealing novel insights from the relationship between molecular measurements and pathology remains a very impactful application of machine learning in biomedicine. Data in this domain typically contain only a few observations but thousands…

Machine Learning · Computer Science 2026-03-31 Christopher Kolberg , Jules Kreuer , Jonas Huurdeman , Sofiane Ouaari , Katharina Eggensperger , Nico Pfeifer

Newsroom in online ecosystem is difficult to untangle. With prevalence of social media, interactions between journalists and individuals become visible, but lack of understanding to inner processing of information feedback loop in public…

Computers and Society · Computer Science 2018-01-03 Pau Perng-Hwa Kung

This paper discusses predictive densities under the Kullback--Leibler loss for high-dimensional Poisson sequence models under sparsity constraints. Sparsity in count data implies zero-inflation. We present a class of Bayes predictive…

Statistics Theory · Mathematics 2020-09-08 Keisuke Yano , Ryoya Kaneko , Fumiyasu Komaki

The concepts of sparsity, and regularised estimation, have proven useful in many high-dimensional statistical applications. Dynamic factor models (DFMs) provide a parsimonious approach to modelling high-dimensional time series, however, it…

Methodology · Statistics 2023-03-22 Luke Mosley , Tak-Shing T. Chan , Alex Gibberd

Differential abundance analysis is at the core of statistical analysis of microbiome data. The compositional nature of microbiome sequencing data makes false positive control challenging. Here, we show that the compositional effects can be…

Methodology · Statistics 2022-03-15 Huijuan Zhou , Kejun He , Jun Chen , Xianyang Zhang

Most of previous works and applications of Bayesian factor model have assumed the normal likelihood regardless of its validity. We propose a Bayesian factor model for heavy-tailed high-dimensional data based on multivariate Student-$t$…

Methodology · Statistics 2020-12-10 Jaejoon Lee , Jaeyong Lee

Within the framework of probability models for overdispersed count data, we propose the generalized fractional Poisson distribution (gfPd), which is a natural generalization of the fractional Poisson distribution (fPd), and the standard…

Probability · Mathematics 2021-01-12 Dexter Cahoy , Elvira Di Nardo , Federico Polito

In this work, a systematic protocol is proposed to automatically parametrize implicit solvent models with polar and nonpolar components. The proposed protocol utilizes the classical Poisson model or the Kohn-Sham density functional theory…

Chemical Physics · Physics 2016-11-03 Bao Wang , Chengzhang Wang , Guowei Wei

Metabolomics is the study of small molecules in biological samples. Metabolomics data are typically high-dimensional and contain highly correlated variables and frequent missing values. Both missing at random (MAR) data, due to acquisition…