English
Related papers

Related papers: Diagnosing overdispersion in longitudinal analyses…

200 papers

We propose a discrete-time, finite-state stationary process that can possess long-range dependence. Among the interesting features of this process is that each state can have different long-term dependency, i.e., the indicator sequence can…

Probability · Mathematics 2022-09-19 Jeonghwa Lee

Classical peaks over threshold analysis is widely used for statistical modeling of sample extremes, and can be supplemented by a model for the sizes of clusters of exceedances. Under mild conditions a compound Poisson process model allows…

Applications · Statistics 2016-08-14 Mária Süveges , Anthony C. Davison

We consider the complex data modeling problem motivated by the zero-inflated and overdispersed data from microbiome studies. Analyzing how microbiome abundance is associated with human biological features, such as BMI, is of great…

Methodology · Statistics 2025-03-31 Zirui Wang , Tianying Wang

Insurance data can be asymmetric with heavy tails, causing inadequate adjustments of the usually applied models. To deal with this issue, hierarchical models for collective risk with heavy-tails of the claims distributions that take also…

Applications · Statistics 2021-01-26 Pamela M. Chiroque-Solano , Fernando A. S. Moura

Logistic regression with unknown sizes has many important applications in biological and medical sciences. All models about this problem in the literature are parametric ones. A semiparametric regression model is proposed. This model…

Statistics Theory · Mathematics 2007-06-13 Wei Zhang

In many longitudinal microarray studies, the gene expression levels in a random sample are observed repeatedly over time under two or more conditions. The resulting time courses are generally very short, high-dimensional, and may have…

Applications · Statistics 2013-02-26 Maurice Berk , Cheryl Hemingway , Michael Levin , Giovanni Montana

Rooted in genetics, human complex diseases are largely influenced by environmental factors. Existing literature has shown the power of integrative gene-environment interaction analysis by considering the joint effect of environmental…

Methodology · Statistics 2022-08-26 Jingyi Zhang , Xu Liu , Honglang Wang , Yuehua Cui

Real-life data are often non-IID due to complex distributions and interactions, and the sensitivity to the distribution of samples can differ among learning models. Accordingly, a key question for any supervised or unsupervised model is…

Machine Learning · Computer Science 2023-10-03 Zhilin Zhao , Longbing Cao

We consider the problem of estimating fold-changes in the expected value of a multivariate outcome observed with unknown sample-specific and category-specific perturbations. This challenge arises in high-throughput sequencing studies of the…

Methodology · Statistics 2026-04-24 David S Clausen , Sarah Teichman , Amy D Willis

Understanding variable dependence, particularly eliciting their statistical properties given a set of covariates, provides the mathematical foundation in practical operations management such as risk analysis and decision-making given…

Methodology · Statistics 2023-09-06 Yunyun Wang , Tatsushi Oka , Dan Zhu

Univariate and multivariate normal probability distributions are widely used when modeling decisions under uncertainty. Computing the performance of such models requires integrating these distributions over specific domains, which can vary…

Machine Learning · Statistics 2024-07-31 Abhranil Das , Wilson S Geisler

The Poisson multinomial distribution (PMD) describes the distribution of the sum of $n$ independent but non-identically distributed random vectors, in which each random vector is of length $m$ with 0/1 valued elements and only one of its…

Computation · Statistics 2022-01-13 Zhengzhi Lin , Yueyao Wang , Yili Hong

Spatio-temporal pathogen spread is often partially observed at the metapopulation scale. Available data correspond to proxies and are incomplete, censored and heterogeneous. Moreover, representing such biological systems often leads to…

Populations and Evolution · Quantitative Biology 2023-12-04 Gaël Beaunée , Pauline Ezanno , Alain Joly , Pierre Nicolas , Elisabeta Vergu

The analysis of count data is commonly done using Poisson models. Negative binomial models are a straightforward and readily motivated generalization for the case of overdispersed data, i.e., when the observed variance is greater than…

Methodology · Statistics 2016-01-06 Christian Röver , Stefan Andreas , Tim Friede

In this paper, we draw attention to a promising yet slightly underestimated measure of variability - the Gini coefficient. We describe two new ways of defining and interpreting this parameter. Using our new representations, we compute the…

Statistics Theory · Mathematics 2022-10-13 Marta Milewska , Remco van der Hofstad , Bert Zwart

We propose a new methodology to detect zero-inflation and overdispersion based on the comparison of the expected sample extremes among convexly ordered distributions. The method is very flexible and includes tests for the proportion of…

Methodology · Statistics 2008-09-25 A. Baillo , J. Carcamo , J. R. Berrendero

RNA-Seq data characteristically exhibits large variances, which need to be appropriately accounted for in the model. We first explore the effects of this variability on the maximum likelihood estimator (MLE) of the overdispersion parameter…

Methodology · Statistics 2015-12-03 Luis Leon-Novelo , Claudio Fuentes , Sarah Emerson

Negative binomial regression is commonly employed to analyze overdispersed count data. With small to moderate sample sizes, the maximum likelihood estimator of the dispersion parameter may be subject to a significant bias, that in turn…

Methodology · Statistics 2020-11-06 Euloge Clovis Kenne Pagui , Alessandra Salvan , Nicola Sartori

Multitype branching processes with immigration in one type are used to model the dynamics of stage-structured plant populations. Parametric inference is first carried out when count data of all types are observed. Statistical…

Applications · Statistics 2009-02-27 Catherine Laredo , Olivier David , Aurélie Garnier

We introduce a new measure of interdependence among the components of a random vector along the main diagonal of the vector copula, i.e. along the line $u_{1}=\ldots=u_{J}$, for $\left(u_{1},\ldots,u_{J}\right)\in\left[0,1\right]^{J}$. Our…

Methodology · Statistics 2014-08-29 Jhan Rodríguez , András Bárdossy