English
Related papers

Related papers: Testing for differential abundance in compositiona…

200 papers

When working with large biological data sets, exploratory analysis is an important first step for understanding the latent structure and for generating hypotheses to be tested in subsequent analyses. However, when the number of variables is…

Methodology · Statistics 2017-02-03 Julia Fukuyama

This study demonstrates the existence of a testable condition for the identification of the causal effect of a treatment on an outcome in observational data, which relies on two sets of variables: observed covariates to be controlled for…

Econometrics · Economics 2026-05-20 Martin Huber , Jannis Kueck

As sequencing technologies become more affordable and genomic databases expand continuously, the reuse of publicly available sequencing data emerges as a powerful strategy for studying microbial pathogens. Indeed, raw sequencing reads…

Quantitative Methods · Quantitative Biology 2025-05-16 Damien Richard , Nils Poulicard

Human microbiome studies use sequencing technologies to measure the abundance of bacterial species or Operational Taxonomic Units (OTUs) in samples of biological material. Typically the data are organized in contingency tables with OTU…

Methodology · Statistics 2017-03-27 Boyu Ren , Sergio Bacallado , Stefano Favaro , Susan Holmes , Lorenzo Trippa

Mutations in a microbial population can increase the frequency of a genotype not only by increasing its exponential growth rate, but also by decreasing its lag time or adjusting the yield (resource efficiency). The contribution of multiple…

Populations and Evolution · Quantitative Biology 2018-02-15 Michael Manhart , Bharat V. Adkar , Eugene I. Shakhnovich

Two species with similar resource requirements respond in a characteristic way to variations in their habitat -- their abundances rise and fall in concert. We use this idea to learn how bacterial populations in the microbiota respond to…

Populations and Evolution · Quantitative Biology 2015-10-02 Charles K. Fisher , Thierry Mora , Aleksandra M. Walczak

We developed a low-cost, high-throughput microbiome profiling method that uses combinatorial sequence tags attached to PCR primers that amplify the rRNA V6 region. Amplified PCR products are sequenced using an Illumina paired-end protocol…

It is well known that correlations in microarray data represent a serious nuisance deteriorating the performance of gene selection procedures. This paper is intended to demonstrate that the correlation structure of microarray data provides…

Applications · Statistics 2007-12-18 Lev Klebanov , Andrei Yakovlev

A number of large spectroscopic surveys of stars in the Milky Way are under way or are being planned. In this context it is important to discuss the extent to which elemental abundances can be used as discriminators between different (known…

Astrophysics of Galaxies · Physics 2015-06-15 Lennart Lindegren , Sofia Feltzing

A frequent challenge encountered with ecological data is how to interpret, analyze, or model data having a high proportion of zeros. Much attention has been given to zero-inflated count data, whereas models for non-negative continuous data…

Methodology · Statistics 2022-05-23 Becky Tang , Henry A Frye , Alan E. Gelfand , John A Silander

Modern generative models exhibit unprecedented capabilities to generate extremely realistic data. However, given the inherent compositionality of the real world, reliable use of these models in practical applications requires that they…

Machine Learning · Computer Science 2025-07-29 Maya Okawa , Ekdeep Singh Lubana , Robert P. Dick , Hidenori Tanaka

The analysis of count data is commonly done using Poisson models. Negative binomial models are a straightforward and readily motivated generalization for the case of overdispersed data, i.e., when the observed variance is greater than…

Methodology · Statistics 2016-01-06 Christian Röver , Stefan Andreas , Tim Friede

Food authenticity studies are concerned with determining if food samples have been correctly labeled or not. Discriminant analysis methods are an integral part of the methodology for food authentication. Motivated by food authenticity…

Methodology · Statistics 2010-10-08 Thomas Brendan Murphy , Nema Dean , Adrian E. Raftery

Recent discoveries suggest that our gut microbiome plays an important role in our health and wellbeing. However, the gut microbiome data are intricate; for example, the microbial diversity in the gut makes the data high-dimensional. While…

Methodology · Statistics 2021-03-02 Fang Xie , Johannes Lederer

Randomized A/B tests within online learning platforms represent an exciting direction in learning sciences. With minimal assumptions, they allow causal effect estimation without confounding bias and exact statistical inference even in small…

Methodology · Statistics 2023-06-13 Adam C. Sales , Ethan B. Prihar , Johann A. Gagnon-Bartsch , Neil T. Heffernan

Count data are common in medical research. When these data have more zeros than expected by the most used count distributions, it is common to employ a zero-inflated regression model. However, the interpretability of these models is much…

Methodology · Statistics 2025-09-30 Gustavo H. A. Pereira , Jeremias Leão , Manoel Santos-Neto , Jianwen Cai

Due to their type of mathematical construction, the use of standard financial ratios in studies analysing the financial health of a group of firms leads to a series of statistical problems that can invalidate the results obtained. These…

Statistical Finance · Quantitative Finance 2025-05-28 Salvador Linares-Mustarós , Maria Àngels Farreras-Noguer , Núria Arimany-Serrat , Germà Coenders

Causal discovery is to learn cause-effect relationships among variables given observational data and is important for many applications. Existing causal discovery methods assume data sufficiency, which may not be the case in many real world…

Machine Learning · Computer Science 2022-06-20 Zijun Cui , Naiyu Yin , Yuru Wang , Qiang Ji

Compositional data arise in many real-life applications and versatile methods for properly analyzing this type of data in the regression context are needed. When parametric assumptions do not hold or are difficult to verify, non-parametric…

Methodology · Statistics 2023-09-07 Michail Tsagris , Abdulaziz Alenazi , Connie Stewart

Missing data are an unavoidable complication frequently encountered in many causal discovery tasks. When a missing process depends on the missing values themselves (known as self-masking missingness), the recovery of the joint distribution…

Machine Learning · Computer Science 2023-12-20 Jie Qiao , Zhengming Chen , Jianhua Yu , Ruichu Cai , Zhifeng Hao