中文
相关论文

相关论文: CAT: a conditional association test for microbiome…

200 篇论文

More than ever, today we are left with the abundance of molecular data outpaced by the advancements of the phylogenomic methods. Especially in the case of presence of many genes over a set of species under the phylogeny question, more…

应用统计 · 统计学 2021-11-29 Ali Amiryousefi

Functional data is a powerful tool for capturing and analyzing complex patterns and relationships in a variety of fields, allowing for more precise modeling, visualization, and decision-making. For example, in healthcare, functional data…

统计方法学 · 统计学 2023-04-26 Xiyuan Gao , Jiayi Wang , Guanyu Hu , Jianguo Sun

Data mining and machine learning techniques such as classification and regression trees (CART) represent a promising alternative to conventional logistic regression for propensity score estimation. Whereas incomplete data preclude the…

机器学习 · 统计学 2018-07-26 Bas B. L. Penning de Vries , Maarten van Smeden , Rolf H. H. Groenwold

We develop a new identification strategy for average treatment effects on the treated (ATT) in panel data with discrete outcomes. Standard difference-in-differences (DiD) relies on parallel trends, which is frequently violated in…

计量经济学 · 经济学 2026-03-10 Young Ahn , Hiroyuki Kasahara

Microbes can affect processes from food production to human health. Such microbes are not isolated, but rather interact with each other and establish connections with their living environments. Understanding these interactions is essential…

应用统计 · 统计学 2021-09-07 Liang Chen , Qiuyan He , Hui Wan , Shun He , Minghua Deng

Conditional independence testing (CIT) is a common task in machine learning, e.g., for variable selection, and a main component of constraint-based causal discovery. While most current CIT approaches assume that all variables are numerical…

机器学习 · 计算机科学 2023-11-07 Oana-Iuliana Popescu , Andreas Gerhardus , Jakob Runge

The objective of clusterability evaluation is to check whether a clustering structure exists within the data set. As a crucial yet often-overlooked issue in cluster analysis, it is essential to conduct such a test before applying any…

机器学习 · 计算机科学 2025-01-07 Lianyu Hu , Junjie Dong , Mudi Jiang , Yan Liu , Zengyou He

Decision tree and random forest classification and regression are some of the most widely used in machine learning approaches. Binary decision tree implementations commonly use conditioning in the form 'feature $\leq$ (or $<$) threshold',…

机器学习 · 计算机科学 2023-12-19 Gábor Timár , György Kovács

Genetic association analyses often involve data from multiple potentially-heterogeneous subgroups. The expected amount of heterogeneity can vary from modest (e.g., a typical meta-analysis) to large (e.g., a strong gene--environment…

统计方法学 · 统计学 2014-04-15 Xiaoquan Wen , Matthew Stephens

Multiple Additive Regression Trees (MART), an ensemble model of boosted regression trees, is known to deliver high prediction accuracy for diverse tasks, and it is widely used in practice. However, it suffers an issue which we call…

机器学习 · 计算机科学 2015-05-11 K. V. Rashmi , Ran Gilad-Bachrach

Robots are often so complex that one person may not know all the ins and outs of the system. Inheriting software and hardware infrastructure with limited documentation and/or practical robot experience presents a costly challenge for an…

机器人学 · 计算机科学 2020-07-24 Victoria Edwards , Loy McGuire , Signe Redfield

Analyzing multivariate count data generated by high-throughput sequencing technology in microbiome research studies is challenging due to the high-dimensional and compositional structure of the data and overdispersion. In practice,…

应用统计 · 统计学 2023-11-03 Jingyan Fu , Matthew D. Koslovsky , Andreas M. Neophytou , Marina Vannucci

There has been a growing acknowledgement of the involvement of the gut microbiome - the collection of microbes that reside in our gut - in regulating our mood and behaviour. This phenomenon is referred to as the microbiome-gut-brain axis.…

基因组学 · 定量生物学 2023-12-11 Thomaz F. S. Bastiaanssen , Thomas P. Quinn , Amy Loughman

Microbiome `omics approaches can reveal intriguing relationships between the human microbiome and certain disease states. Along with the identification of specific bacteria taxa associated with diseases, recent scientific advancements…

应用统计 · 统计学 2019-10-07 Shuang Jiang , Guanghua Xiao , Andrew Y. Koh , Qiwei Li , Xiaowei Zhan

Recent multi-omic microbiome studies enable integrative analysis of microbes and metabolites, uncovering their associations with various host conditions. Such analyses require multivariate models capable of accounting for the complex…

Microorganisms play critical roles in human health and disease. It is well known that microbes live in diverse communities in which they interact synergistically or antagonistically. Thus for estimating microbial associations with clinical…

There exist a wide range of single number metrics for assessing performance of classification algorithms, including AUC and the F1-score (Wikipedia lists 17 such metrics, with 27 different names). In this article, I propose a new metric to…

机器学习 · 计算机科学 2023-11-21 David J. T. Sumpter

Longitudinal data are common in clinical trials and observational studies, where missing outcomes due to dropouts are always encountered. Under such context with the assumption of missing at random, the weighted generalized estimating…

统计方法学 · 统计学 2019-04-30 Chixiang Chen , Biyi Shen , Lijun Zhang , Yuan Xue , Ming Wang

For the outcomes and phenotypes of complex diseases, multiple types of molecular (genetic, genomic, epigenetic, etc.) changes, environmental risk factors, and their interactions have been found to have important contributions. In each of…

统计方法学 · 统计学 2019-12-19 Yaqing Xu , Mengyun Wu , Shuangge Ma

In human microbiome studies, sequencing reads data are often summarized as counts of bacterial taxa at various taxonomic levels specified by a taxonomic tree. This paper considers the problem of analyzing two repeated measurements of…

应用统计 · 统计学 2017-02-17 Pixu Shi , Hongzhe Li