中文
相关论文

相关论文: LinDA: linear models for differential abundance an…

200 篇论文

Analyzing multivariate count data generated by high-throughput sequencing technology in microbiome research studies is challenging due to the high-dimensional and compositional structure of the data and overdispersion. In practice,…

应用统计 · 统计学 2023-11-03 Jingyan Fu , Matthew D. Koslovsky , Andreas M. Neophytou , Marina Vannucci

Biological sequencing data consist of read counts, e.g. of specified taxa and often exhibit sparsity (zero-count inflation) and overdispersion (extra-Poisson variability). As most sequencing techniques provide an arbitrary total count,…

应用统计 · 统计学 2024-07-01 Noora Kartiosuo , Jaakko Nevalainen , Olli Raitakari , Katja Pahkala , Kari Auranen

Compositional data sets are ubiquitous in science, including geology, ecology, and microbiology. In microbiome research, compositional data primarily arise from high-throughput sequence-based profiling experiments. These data comprise…

统计理论 · 数学 2019-03-05 Patrick L. Combettes , Christian L. Müller

The focus of this paper is to extend Fisher's linear discriminant analysis (LDA) to both densely re-corded functional data and sparsely observed longitudinal data for general $c$-category classification problems. We propose an efficient…

统计方法学 · 统计学 2017-09-07 Lu-Hung Chen , Ci-Ren Jiang

High-dimensional compositional data, such as those from human microbiome studies, pose unique statistical challenges due to the simplex constraint and excess zeros. While dimension reduction is indispensable for analyzing such data,…

统计方法学 · 统计学 2025-09-09 Junyoung Park , Cheolwoo Park , Jeongyoun Ahn

This paper studies Difference-in-Differences (DiD) setups with repeated cross-sectional data and potential compositional changes across time periods. We begin our analysis by deriving the efficient influence function and the semiparametric…

计量经济学 · 经济学 2025-11-17 Pedro H. C. Sant'Anna , Qi Xu

The human microbiome plays critical roles in human health and has been linked to many diseases. While advanced sequencing technologies can characterize the composition of the microbiome in unprecedented detail, it remains challenging to…

Using a sample from a population to estimate the proportion of the population with a certain category label is a broadly important problem. In the context of microbiome studies, this problem arises when researchers wish to use a sample from…

统计方法学 · 统计学 2019-02-08 Bryan D. Martin , Daniela Witten , Amy D. Willis

Data augmentation plays a key role in modern machine learning pipelines. While numerous augmentation strategies have been studied in the context of computer vision and natural language processing, less is known for other data modalities.…

机器学习 · 统计学 2022-05-23 Elliott Gordon-Rodriguez , Thomas P. Quinn , John P. Cunningham

There has been a growing acknowledgement of the involvement of the gut microbiome - the collection of microbes that reside in our gut - in regulating our mood and behaviour. This phenomenon is referred to as the microbiome-gut-brain axis.…

基因组学 · 定量生物学 2023-12-11 Thomaz F. S. Bastiaanssen , Thomas P. Quinn , Amy Loughman

Linear Discriminant Analysis (LDA) is a fundamental method for classification. Its simple linear structure facilitates interpretation, and it is naturally suited to multi-class settings. LDA is also closely connected to several classical…

统计方法学 · 统计学 2026-04-09 Xin Bing , Bingqing Li , Marten Wegkamp

Existing effect measures for compositional features are inadequate for many modern applications, for example, in microbiome research, since they display traits such as high-dimensionality and sparsity that can be poorly modelled with…

统计方法学 · 统计学 2025-06-02 Anton Rask Lundborg , Niklas Pfister

Compositional data have two unique characteristics compared to typical multivariate data: the observed values are nonnegative and their summand is exactly one. To reflect these characteristics, a specific regularized regression model with…

机器学习 · 计算机科学 2018-12-24 Jong-June Jeon , Yongdai Kim , Sungho Won , Hosik Choi

High-dimensional data must be highly structured to be learnable. Although the compositional and hierarchical nature of data is often put forward to explain learnability, quantitative measurements establishing these properties are scarce.…

机器学习 · 统计学 2025-03-04 Antonio Sclocchi , Alessandro Favero , Noam Itzhak Levi , Matthieu Wyart

With the development of next generation sequencing technology, researchers have now been able to study the microbiome composition using direct sequencing, whose output are bacterial taxa counts for each microbiome sample. One goal of…

应用统计 · 统计学 2013-05-24 Jun Chen , Hongzhe Li

Microbiome `omics approaches can reveal intriguing relationships between the human microbiome and certain disease states. Along with the identification of specific bacteria taxa associated with diseases, recent scientific advancements…

应用统计 · 统计学 2019-10-07 Shuang Jiang , Guanghua Xiao , Andrew Y. Koh , Qiwei Li , Xiaowei Zhan

Compositional data arise in many areas of research in the natural and biomedical sciences. One prominent example is in the study of the human gut microbiome, where one can measure the relative abundance of many distinct microorganisms in a…

统计方法学 · 统计学 2024-04-26 Aaron J. Molstad , Karl Oskar Ekvall , Piotr M. Suder

This paper introduces a rectified and renormalized Fisher-Bingham model for compositional data with zeros, motivated in part by the presence of zeros in microbiota studies. The approach represents compositions through a square-root…

统计方法学 · 统计学 2026-04-29 Eugene Han , Marahi Perez-Tamayo , Hannah D. Holscher , Ruoqing Zhu

Inferring causality using longitudinal observational databases is challenging due to the passive way the data are collected. The majority of associations found within longitudinal observational data are often non-causal and occur due to…

计算工程、金融与科学 · 计算机科学 2016-11-17 Jenna Reps , Uwe Aickelin

The growing use of high-throughput sequencing (HTS) has enabled the large-scale production of compositional count data, driving progress in microbiome research. However, such count data are often high-dimensional, over-dispersed, and…

其他统计学 · 统计学 2026-05-22 Wenqi Tang , Kamila Fačevicová , Klaus Nordhausen , Sara Taskinen