English
Related papers

Related papers: Testing for differential abundance in compositiona…

200 papers

In microbiome studies, it is of interest to use a sample from a population of microbes, such as the gut microbiota community, to estimate the population proportion of these taxa. However, due to biases introduced in sampling and…

Methodology · Statistics 2022-10-11 Roulan Jiang , Xiang Zhan , Tianying Wang

A frequent challenge encountered with compositional ecological data is how to interpret and model data with a high proportion of zeros and $N$'s. Such data frequently occur in ecological applications where counts of species are collected…

Methodology · Statistics 2025-08-04 James Sweeney , John Haslett , Dipankar Bandyopadhyay , Michael Fop , Andrew C. Parnell

The human microbiome can contribute to pathogeneses of many complex diseases by mediating disease-leading causal pathways. However, standard mediation analysis methods are not adequate to analyze the microbiome as a mediator due to the…

Existing effect measures for compositional features are inadequate for many modern applications, for example, in microbiome research, since they display traits such as high-dimensionality and sparsity that can be poorly modelled with…

Methodology · Statistics 2025-06-02 Anton Rask Lundborg , Niklas Pfister

We propose a simple data augmentation protocol aimed at providing a compositional inductive bias in conditional and unconditional sequence models. Under this protocol, synthetic training examples are constructed by taking real training…

Computation and Language · Computer Science 2020-05-20 Jacob Andreas

A key challenge in differential abundance analysis of microbial samples is that the counts for each sample are compositional, resulting in biased comparisons of the absolute abundance across study groups. Normalization-based differential…

Genomics · Quantitative Biology 2024-11-26 Dylan Clark-Boucher , Brent A Coull , Harrison T Reeder , Fenglei Wang , Qi Sun , Jacqueline R Starr , Kyu Ha Lee

The Dirichlet-multinomial (DM) distribution plays a fundamental role in modern statistical methodology development and application. Recently, the DM distribution and its variants have been used extensively to model multivariate count data…

Methodology · Statistics 2023-02-27 Matthew D. Koslovsky

In compositional data, an observation is a vector with non-negative components which sum to a constant, typically 1. Data of this type arise in many areas, such as geology, archaeology, biology, economics and political science among others.…

Methodology · Statistics 2015-06-18 Michail Tsagris

Recent evidence suggests that analyzing the presence/absence of taxonomic features can offer a compelling alternative to differential abundance analysis in microbiome studies. However, standard approaches to differential prevalence analysis…

Methodology · Statistics 2026-05-26 Juho Pelto , Kari Auranen , Janne V. Kujala , Leo Lahti

Compositional data (i.e., data comprising random variables that sum up to a constant) arises in many applications including microbiome studies, chemical ecology, political science, and experimental designs. Yet when compositional data serve…

Methodology · Statistics 2025-01-03 Ritwik Bhaduri , Siyuan Ma , Lucas Janson

The advent of high-throughput sequencing technologies has made data from DNA material readily available, leading to a surge of microbiome-related research establishing links between markers of microbiome health and specific outcomes.…

Methodology · Statistics 2018-06-19 Susheela P. Singh , Ana-Maria Staicu , Robert R. Dunn , Noah Fierer , Brian J. Reich

Quantification of microbial interactions from 16S rRNA and meta-genomic sequencing data is difficult due to their sparse nature, as well as the fact that the data only provides measures of relative abundance. In this paper, we propose using…

Methodology · Statistics 2021-11-04 Rebecca A. Deek , Hongzhe Li

Data augmentation plays a key role in modern machine learning pipelines. While numerous augmentation strategies have been studied in the context of computer vision and natural language processing, less is known for other data modalities.…

Machine Learning · Statistics 2022-05-23 Elliott Gordon-Rodriguez , Thomas P. Quinn , John P. Cunningham

Microbial ecosystems are remarkably diverse, stable, and often consist of a balanced mixture of core and peripheral species. Here we propose a conceptual model exhibiting all these emergent properties in quantitative agreement with real…

Populations and Evolution · Quantitative Biology 2018-04-18 Akshit Goyal , Sergei Maslov

We develop statistically based methods to detect single nucleotide DNA mutations in next generation sequencing data. Sequencing generates counts of the number of times each base was observed at hundreds of thousands to billions of genome…

Applications · Statistics 2012-10-01 Omkar Muralidharan , Georges Natsoulis , John Bell , Hanlee Ji , Nancy R. Zhang

We introduce a novel approach to compositional data analysis based on $L^{\infty}$-normalization, addressing challenges posed by zero-rich high-throughput data. Traditional methods like Aitchison's transformations require excluding zeros,…

Computation · Statistics 2025-03-28 Pawel Gajer , Jacques Ravel

We argue that the selective inclusion of data points based on latent objectives is common in practical situations, such as music sequences. Since this selection process often distorts statistical analysis, previous work primarily views it…

Machine Learning · Computer Science 2024-07-02 Yujia Zheng , Zeyu Tang , Yiwen Qiu , Bernhard Schölkopf , Kun Zhang

Testing differences in mean vectors is a fundamental task in the analysis of high-dimensional compositional data. Existing methods may suffer from low power if the underlying signal pattern is in a situation that does not favor the deployed…

Methodology · Statistics 2025-03-11 Danning Li , Lingzhou Xue , Haoyi Yang , Xiufan Yu

Anomaly detection when observing a large number of data streams is essential in a variety of applications, ranging from epidemiological studies to monitoring of complex systems. High-dimensional scenarios are usually tackled with…

Methodology · Statistics 2025-12-18 Ivo V. Stoepker , Rui M. Castro , Ery Arias-Castro , Edwin van den Heuvel

Advances in data collection are producing growing volumes of temporal count observations, making adapted modeling increasingly necessary. In this work, we introduce a generative framework for independent component analysis of temporal count…

Methodology · Statistics 2026-01-30 Alexandre Chaussard , Anna Bonnet , Sylvain Le Corff