English
Related papers

Related papers: A critical comparison of handling zeros in high-di…

200 papers

Microbiome compositional data are often high-dimensional, sparse, and exhibit pervasive cross-sample heterogeneity. Generative modeling is a popular approach to analyze such data, and effective generative models must accurately characterize…

Methodology · Statistics 2025-01-03 Zhuoqun Wang , Jialiang Mao , Li Ma

High-dimensional compositional data are prevalent in many applications. The simplex constraint poses intrinsic challenges to inferring the conditional dependence relationships among the components forming a composition, as encoded by a…

Methodology · Statistics 2024-03-25 Shucong Zhang , Huiyuan Wang , Wei Lin

Compositional data, also referred to as simplicial data, naturally arise in many scientific domains such as geochemistry, microbiology, and economics. In such domains, obtaining sensible lower-dimensional representations and modes of…

Count data are common in medical research. When these data have more zeros than expected by the most used count distributions, it is common to employ a zero-inflated regression model. However, the interpretability of these models is much…

Methodology · Statistics 2025-09-30 Gustavo H. A. Pereira , Jeremias Leão , Manoel Santos-Neto , Jianwen Cai

High-dimensional compositional data are frequently encountered in many fields of modern scientific research. In regression analysis of compositional data, the presence of covariate measurement errors poses grand challenges for existing…

Methodology · Statistics 2024-07-23 Wenxi Tan , Lingzhou Xue , Songshan Yang , Xiang Zhan

Differential abundance analysis is at the core of statistical analysis of microbiome data. The compositional nature of microbiome sequencing data makes false positive control challenging. Here, we show that the compositional effects can be…

Methodology · Statistics 2022-03-15 Huijuan Zhou , Kejun He , Jun Chen , Xianyang Zhang

Differential abundance tests in compositional data are essential and fundamental tasks in various biomedical applications, such as single-cell, bulk RNA-seq, and microbiome data analysis. However, because of the compositional constraint and…

Methodology · Statistics 2022-04-14 Shulei Wang

The human microbiome can contribute to pathogeneses of many complex diseases by mediating disease-leading causal pathways. However, standard mediation analysis methods are not adequate to analyze the microbiome as a mediator due to the…

Objective: The proper handling of missing values is critical to delivering reliable estimates and decisions, especially in high-stakes fields such as clinical research. The increasing diversity and complexity of data have led many…

We consider the analysis of count data in which the observed frequency of zero counts is unusually large, typically with respect to the Poisson distribution. We focus on two alternative modelling approaches: Over-Dispersion (OD) models, and…

Methodology · Statistics 2021-07-30 John Haslett , Andrew C. Parnell , John Hinde , Rafael A. Moral

In compositional data analysis an observation is a vector containing non-negative values, only the relative sizes of which are considered to be of interest. Without loss of generality, a compositional vector can be taken to be a vector of…

Methodology · Statistics 2015-06-18 Michail Tsagris , Simon Preston , Andrew T. A. Wood

This paper proposes a new generalized linear model with the fractional binomial distribution. Zero-inflated Poisson/negative binomial distributions are used for count data with many zeros. To analyze the association of such a count variable…

Methodology · Statistics 2025-08-01 Jeonghwa Lee , Chloe Breece

The study of immune cellular composition has been of great scientific interest in immunology because of the generation of multiple large-scale data. From the statistical point of view, such immune cellular data should be treated as…

Applications · Statistics 2022-04-22 Jinkyung Yoo , Zequn Sun , Michael Greenacre , Qin Ma , Dongjun Chung , Young Min Kim

High throughput sequencing is a technology that allows for the generation of millions of reads of genomic data regarding a study of interest, and data from high throughput sequencing platforms are usually count compositions. Subsequent…

Quantitative Methods · Quantitative Biology 2017-04-07 Jia R. Wu , Jean M. Macklaim , Briana L. Genge , Gregory B. Gloor

Due to their type of mathematical construction, the use of standard financial ratios in studies analysing the financial health of a group of firms leads to a series of statistical problems that can invalidate the results obtained. These…

Statistical Finance · Quantitative Finance 2025-05-28 Salvador Linares-Mustarós , Maria Àngels Farreras-Noguer , Núria Arimany-Serrat , Germà Coenders

Large-scale classification of data where classes are structurally organized in a hierarchy is an important area of research. Top-down approaches that exploit the hierarchy during the learning and prediction phase are efficient for large…

Machine Learning · Computer Science 2017-06-06 Azad Naik , Huzefa Rangwala

Large biological datasets are being produced at a rapid pace and create substantial storage challenges, particularly in the domain of high-throughput sequencing (HTS). Most approaches currently used to store HTS data are either unable to…

Quantitative Methods · Quantitative Biology 2014-03-05 Fabien Campagne , Kevin C. Dorff , Nyasha Chambwe , James T. Robinson , Jill P. Mesirov , Thomas D. Wu

Recent approaches in music generation rely on disentangled representations, often labeled as structure and timbre or local and global, to enable controllable synthesis. Yet the underlying properties of these embeddings remain underexplored.…

Data corruption, including missing and noisy data, poses significant challenges in real-world machine learning. This study investigates the effects of data corruption on model performance and explores strategies to mitigate these effects…

Machine Learning · Computer Science 2025-05-22 Qi Liu , Wanjing Ma

Mortality forecasting is crucial for demographic planning and actuarial studies, especially for projecting population ageing and longevity risk. Classical approaches largely rely on extrapolative methods, such as the Lee-Carter (LC) model,…

Applications · Statistics 2026-02-24 Han Ying Lim , Dharini Pathmanathan , Sophie Dabo-Niang