English
Related papers

Related papers: Modeling data with zero inflation and overdispersi…

200 papers

The main object of this article is to present an extension of the zero-inflated Poisson-Lindley distribution, called of zero-modified Poisson-Lindley. The additional parameter $\pi$ of the zero-modified Poisson-Lindley has a natural…

Methodology · Statistics 2018-11-28 Danillo Xavier , Manoel Santos-Neto , Marcelo Bourguignon , Vera Tomazella

The popular generalized additive model framework is extended to allow both the mean curves and the response distribution to be nonparametric. The approach is demonstrated to be a flexible yet parsimonious tool for data analysis in its own…

Methodology · Statistics 2017-09-18 Alan Huang , Nanxi Zhang

Overparameterized models fail to generalize well in the presence of data imbalance even when combined with traditional techniques for mitigating imbalances. This paper focuses on imbalanced classification datasets, in which a small subset…

Machine Learning · Computer Science 2022-06-28 Tina Behnia , Ke Wang , Christos Thrampoulidis

Analyzing overdispersed, zero-inflated, longitudinal count data poses significant modeling and computational challenges, which standard count models (e.g., Poisson or negative binomial mixed effects models) fail to adequately address. We…

Methodology · Statistics 2026-02-11 John Barrera , Ana Arribas-Gil , Dae-Jin Lee , Cristian Meza

In order to better fit real-world datasets, studying asymmetric distribution is of great interest. In this work, we derive several mathematical properties of a general class of asymmetric distributions with positive support which shows up…

Statistics Theory · Mathematics 2025-12-11 Felipe S. Quintino , Pushpa N. Rathie , Luan C. S. M. Ozelim , Tiago A. da Fonseca , Roberto Vila

The M5 competition uncertainty track aims for probabilistic forecasting of sales of thousands of Walmart retail goods. We show that the M5 competition data faces strong overdispersion and sporadic demand, especially zero demand. We discuss…

Machine Learning · Statistics 2021-11-11 Florian Ziel

Latent space models are powerful statistical tools for modeling and understanding network data. While the importance of accounting for uncertainty in network analysis has been well recognized, the current literature predominantly focuses on…

Statistics Theory · Mathematics 2025-08-15 Jinming Li , Shihao Wu , Chengyu Cui , Gongjun Xu , Ji Zhu

Copulas, generalized estimating equations, and generalized linear mixed models promote the analysis of grouped data where non-normal responses are correlated. Unfortunately, parameter estimation remains challenging in these three…

Methodology · Statistics 2024-10-16 Sarah S. Ji , Benjamin B. Chu , Hua Zhou , Kenneth Lange

Imbalanced classification and spurious correlation are common challenges in data science and machine learning. Both issues are linked to data imbalance, with certain groups of data samples significantly underrepresented, which in turn would…

Machine Learning · Statistics 2026-02-10 Ryumei Nakada , Yichen Xu , Lexin Li , Linjun Zhang

Non-gaussian spatial data are very common in many disciplines. For instance, count data are common in disease mapping, and binary data are common in ecology. When fitting spatial regressions for such data, one needs to account for…

Methodology · Statistics 2010-12-01 John Hughes , Murali Haran

Univariate regression models have rich literature for counting data. However, this is not the case for multivariate count data. Therefore, we present the Multivariate Generalized Linear Mixed Models framework that deals with a multivariate…

Count-compositional data arise in many different fields, including high-throughput sequencing experiments, ecological surveys, and palaeoclimate studies, where a common, important goal is to understand how covariates relate to the observed…

Methodology · Statistics 2026-04-10 André F. B. Menezes , Andrew C. Parnell , Keefe Murphy

Motivated by distributed machine learning settings such as Federated Learning, we consider the problem of fitting a statistical model across a distributed collection of heterogeneous data sets whose similarity structure is encoded by a…

Statistics Theory · Mathematics 2021-11-30 Dominic Richards , Sahand N. Negahban , Patrick Rebeschini

Several statistical models used in genome-wide prediction assume independence of marker allele substitution effects, but it is known that these effects might be correlated. In statistics, graphical models have been identified as a useful…

Quantitative Methods · Quantitative Biology 2017-04-13 Carlos Alberto Martínez , Kshitij Khare , Syed Rahman , Mauricio A. Elzo

The assumption of Gaussian or Gaussian mixture data has been extensively exploited in a long series of precise performance analyses of machine learning (ML) methods, on large datasets having comparably numerous samples and features. To…

Machine Learning · Statistics 2025-03-14 Xiaoyi Mai , Zhenyu Liao

Logistic regression with unknown sizes has many important applications in biological and medical sciences. All models about this problem in the literature are parametric ones. A semiparametric regression model is proposed. This model…

Statistics Theory · Mathematics 2007-06-13 Wei Zhang

1. Joint species distribution models (JSDMs) have gained considerable traction among ecologists over the past decade, due to their capacity to answer a wide range of questions at both the species- and the community-level. The family of…

Methodology · Statistics 2024-03-19 Pekka Korhonen , Francis K. C. Hui , Jenni Niku , Sara Taskinen , Bert van der Veen

Missing data arises when certain values are not recorded or observed for variables of interest. However, most of the statistical theory assume complete data availability. To address incomplete databases, one approach is to fill the gaps…

Factor analysis for high-dimensional data is a canonical problem in statistics and has a wide range of applications. However, there is currently no factor model tailored to effectively analyze high-dimensional count responses with…

Methodology · Statistics 2024-08-21 Wei Liu , Qingzhi Zhong

Generalized linear models (GLMs) have been used quite effectively in the modeling of a mean response under nonstandard conditions, where discrete as well as continuous data distributions can be accommodated. The choice of design for a GLM…

Statistics Theory · Mathematics 2016-08-14 André I. Khuri , Bhramar Mukherjee , Bikas K. Sinha , Malay Ghosh