English
Related papers

Related papers: Variable selection for sparse Dirichlet-multinomia…

200 papers

An important task in microbiome studies is to test the existence of and give characterization to differences in the microbiome composition across groups of samples. Important challenges of this problem include the large within-group…

Methodology · Statistics 2019-05-07 Jialiang Mao , Yuhan Chen , Li Ma

When analyzing data from multiple sources, it is often convenient to strike a careful balance between two goals: capturing the heterogeneity of the samples and sharing information across them. We introduce a novel framework to model a…

Methodology · Statistics 2026-03-02 Laura D'Angelo , Bernardo Nipoti , Andrea Ongaro

Structured additive distributional copula regression allows to model the joint distribution of multivariate outcomes by relating all distribution parameters to covariates. Estimation via statistical boosting enables accounting for…

We consider the problem of sparse variable selection in nonparametric additive models, with the prior knowledge of the structure among the covariates to encourage those variables within a group to be selected jointly. Previous works either…

Machine Learning · Computer Science 2012-06-22 Junming Yin , Xi Chen , Eric Xing

Recent advances in bioinformatics have made high-throughput microbiome data widely available, and new statistical tools are required to maximize the information gained from these data. For example, analysis of high-dimensional microbiome…

Methodology · Statistics 2017-03-23 Neal S. Grantham , Brian J. Reich , Elizabeth T. Borer , Kevin Gross

Contamination can severely distort an estimator unless the estimation procedure is suitably robust. This is a well-known issue and has been addressed in Robust Statistics, however, the relation of contamination and distorted variable…

Statistics Theory · Mathematics 2022-07-15 Tino Werner

Biological sequencing data consist of read counts, e.g. of specified taxa and often exhibit sparsity (zero-count inflation) and overdispersion (extra-Poisson variability). As most sequencing techniques provide an arbitrary total count,…

Applications · Statistics 2024-07-01 Noora Kartiosuo , Jaakko Nevalainen , Olli Raitakari , Katja Pahkala , Kari Auranen

We examine the linear regression problem in a challenging high-dimensional setting with correlated predictors where the vector of coefficients can vary from sparse to dense. In this setting, we propose a combination of probabilistic…

Methodology · Statistics 2025-05-13 Roman Parzer , Peter Filzmoser , Laura Vana-Gür

Motivated by the challenges in analyzing gut microbiome and metagenomic data, this work aims to tackle the issue of measurement errors in high-dimensional regression models that involve compositional covariates. This paper marks a…

Methodology · Statistics 2024-09-13 Huali Zhao , Tianying Wang

By creating networks of biochemical pathways, communities of micro-organisms are able to modulate the properties of their environment and even the metabolic processes within their hosts. Next-generation high-throughput sequencing has led to…

Applications · Statistics 2023-03-28 Molly G. Hayes , Morgan G. I. Langille , Hong Gu

Identifying which taxa in our microbiota are associated with traits of interest is important for advancing science and health. However, the identification is challenging because the measured vector of taxa counts (by amplicon sequencing) is…

Genomics · Quantitative Biology 2020-03-31 Barak Brill , Amnon Amir , Ruth Heller

Identifying relevant factors that influence the multinomial counts in compositional data is difficult in high dimensional settings due to the complex associations and overdispersion. Multivariate count models such as the…

Methodology · Statistics 2025-04-08 Alysha Cooper , Zeny Feng , Ayesha Ali , Tim Arciszewski , Lorna Deeth

Food authenticity studies are concerned with determining if food samples have been correctly labeled or not. Discriminant analysis methods are an integral part of the methodology for food authentication. Motivated by food authenticity…

Methodology · Statistics 2010-10-08 Thomas Brendan Murphy , Nema Dean , Adrian E. Raftery

Metagenomics sequencing is routinely applied to quantify bacterial abundances in microbiome studies, where the bacterial composition is estimated based on the sequencing read counts. Due to limited sequencing depth and DNA dropouts, many…

Methodology · Statistics 2019-04-26 Yuanpei Cao , Anru Zhang , Hongzhe Li

Quantile regression is useful for characterizing the conditional distribution of a response variable and understanding heterogeneity in the covariate effects at different quantiles. The rise of high-dimensional physiological data in…

Methodology · Statistics 2026-03-25 Yuanzhen Yue , Stella Self , Yichao Wu , Jiajia Zhang , Rahul Ghosal

Learning governing equations from a family of data sets which share the same physical laws but differ in bifurcation parameters is challenging. This is due, in part, to the wide range of phenomena that could be represented in the data sets…

Numerical Analysis · Mathematics 2017-09-07 Hayden Schaeffer , Giang Tran , Rachel Ward

Dynamic treatment regimes (DTRs) consist of a sequence of decision rules, one per stage of intervention, that finds effective treatments for individual patients according to patient information history. DTRs can be estimated from models…

Methodology · Statistics 2021-12-07 Zeyu Bian , Erica EM Moodie , Susan M Shortreed , Sahir Bhatnagar

Diffusion-based generative models (DBGMs) perturb data to a target noise distribution and reverse this process to generate samples. The choice of noising process, or inference diffusion process, affects both likelihoods and sample quality.…

Machine Learning · Computer Science 2023-03-06 Raghav Singhal , Mark Goldstein , Rajesh Ranganath

The model interpretation is essential in many application scenarios and to build a classification model with a ease of model interpretation may provide useful information for further studies and improvement. It is common to encounter with a…

Machine Learning · Statistics 2019-01-07 Wan-Ping Nicole Chen , Yuan-chin Ivan Chang

The Regression Discontinuity Design (RDD) is a quasi-experimental design that estimates the causal effect of a treatment when its assignment is defined by a threshold value for a continuous assignment variable. The RDD assumes that subjects…

Applications · Statistics 2020-03-27 Federico Ricciardi , Silvia Liverani , Gianluca Baio