English
Related papers

Related papers: A Bayesian Nonparametric Approach for Identifying …

200 papers

Microbiome data analysis is essential for understanding host health and disease, yet its inherent sparsity and noise pose major challenges for accurate imputation, hindering downstream tasks such as biomarker discovery. Existing imputation…

Machine Learning · Computer Science 2025-08-01 Rabeya Tus Sadia , Qiang Cheng

This paper proposes a hierarchical Bayesian multitask learning model that is applicable to the general multi-task binary classification learning problem where the model assumes a shared sparsity structure across different tasks. We derive a…

Many clinical endpoint measures, such as the number of standard drinks consumed per week or the number of days that patients stayed in the hospital, are count data with excessive zeros. However, the zero-inflated nature of such outcomes is…

Applications · Statistics 2022-07-14 Zhengyang Zhou , Minge Xie , David Huh , Eun-Young Mun

Applications of high-dimensional regression often involve multiple sources or types of covariates. We propose methodology for this setting, emphasizing the "wide data" regime with large total dimensionality p and sample size n<<p. We focus…

An important task in microbiome studies is to test the existence of and give characterization to differences in the microbiome composition across groups of samples. Important challenges of this problem include the large within-group…

Methodology · Statistics 2019-05-07 Jialiang Mao , Yuhan Chen , Li Ma

Recent advances in bioinformatics have made high-throughput microbiome data widely available, and new statistical tools are required to maximize the information gained from these data. For example, analysis of high-dimensional microbiome…

Methodology · Statistics 2017-03-23 Neal S. Grantham , Brian J. Reich , Elizabeth T. Borer , Kevin Gross

In human microbiome studies, sequencing reads data are often summarized as counts of bacterial taxa at various taxonomic levels specified by a taxonomic tree. This paper considers the problem of analyzing two repeated measurements of…

Applications · Statistics 2017-02-17 Pixu Shi , Hongzhe Li

It is now practically the norm for data to be very high dimensional in areas such as genetics, machine vision, image analysis and many others. When analyzing such data, parametric models are often too inflexible while nonparametric…

Methodology · Statistics 2011-05-31 Abhishek Bhattacharya , Garritt Page , David Dunson

Differential abundance analysis is a key component of microbiome studies. Although dozens of methods exist there is currently no consensus on the preferred methods. While the correctness of results in differential abundance analysis is an…

Applications · Statistics 2025-04-01 Juho Pelto , Kari Auranen , Janne Kujala , Leo Lahti

The rapid generation of complex, highly skewed, and zero-inflated multi-source count data poses significant challenges for variable selection, particularly in biomedical domains like tumor development and metabolic dysregulation. To address…

Applications · Statistics 2025-11-11 Shan Tang , Shanjun Mao , Shourong Ma , Falong Tan

Multimodal clinical data are characterized by high dimensionality, heterogeneous representations, and structured missingness, posing significant challenges for predictive modeling, data integration, and interpretability. We propose BIONIC…

Microbial communities play important roles in the function and maintenance of various biosystems, ranging from human body to the environment. Current methods for analysis of microbial communities are typically based on taxonomic…

Genomics · Quantitative Biology 2015-12-02 Ehsaneddin Asgari , Kiavash Garakani , Mohammad R. K Mofrad

Datasets with hundreds of variables and many missing values are commonplace. In this setting, it is both statistically and computationally challenging to detect true predictive relationships between variables and also to suppress false…

Machine Learning · Statistics 2018-04-03 Feras Saad , Vikash Mansinghka

The goal of this presentation is to build an efficient non-parametric Bayes classifier in the presence of large numbers of predictors. When analyzing such data, parametric models are often too inflexible while non-parametric procedures tend…

Methodology · Statistics 2013-01-07 Abhishek Bhattacharya

Multi-category data arise in diverse fields including marketing, chemistry, public policy, genomics, political science, and ecology. We consider the problem of estimating ratios of category-specific means in a fully nonparametric setting,…

Methodology · Statistics 2025-10-29 Grant Hopkins , Sarah Teichman , Ellen Graham , Amy D Willis

Inferring microbial interaction networks from abundance patterns is an important approach to advance our understanding of microbial communities in general and the human microbiome in particular. Here we suggest discriminating two levels of…

Populations and Evolution · Quantitative Biology 2023-06-06 Isabella-Hilda Mendler , Barbara Drossel , Marc-Thorsten Hütt

The assessment of diversity and similarity is relevant in monitoring the status of ecosystems. The respective indicators are based on the taxonomic composition of biological communities of interest, currently estimated through the…

Applications · Statistics 2018-10-12 Fabio Divino , Johanna Ärje , Antti Penttinen , Kristian Meissner , Salme Kärkkäinen

In recent years, conditional copulas, that allow dependence between variables to vary according to the values of one or more covariates, have attracted increasing attention. In high dimension, vine copulas offer greater flexibility compared…

Methodology · Statistics 2021-09-24 Rosario Barone , Luciana Dalla Valle

Bayesian nonparametric methods are a popular choice for analysing survival data due to their ability to flexibly model the distribution of survival times. These methods typically employ a nonparametric prior on the survival function that is…

Methodology · Statistics 2022-02-22 Edwin Fong , Brieuc Lehmann

Discrete data are abundant and often arise as counts or rounded data. These data commonly exhibit complex distributional features such as zero-inflation, over-/under-dispersion, boundedness, and heaping, which render many parametric models…

Methodology · Statistics 2023-02-27 Daniel R. Kowal , Bohan Wu