中文
相关论文

相关论文: Variable selection for sparse Dirichlet-multinomia…

200 篇论文

An important task in microbiome studies is to test the existence of and give characterization to differences in the microbiome composition across groups of samples. Important challenges of this problem include the large within-group…

统计方法学 · 统计学 2019-05-07 Jialiang Mao , Yuhan Chen , Li Ma

When analyzing data from multiple sources, it is often convenient to strike a careful balance between two goals: capturing the heterogeneity of the samples and sharing information across them. We introduce a novel framework to model a…

统计方法学 · 统计学 2026-03-02 Laura D'Angelo , Bernardo Nipoti , Andrea Ongaro

Structured additive distributional copula regression allows to model the joint distribution of multivariate outcomes by relating all distribution parameters to covariates. Estimation via statistical boosting enables accounting for…

We consider the problem of sparse variable selection in nonparametric additive models, with the prior knowledge of the structure among the covariates to encourage those variables within a group to be selected jointly. Previous works either…

机器学习 · 计算机科学 2012-06-22 Junming Yin , Xi Chen , Eric Xing

Recent advances in bioinformatics have made high-throughput microbiome data widely available, and new statistical tools are required to maximize the information gained from these data. For example, analysis of high-dimensional microbiome…

统计方法学 · 统计学 2017-03-23 Neal S. Grantham , Brian J. Reich , Elizabeth T. Borer , Kevin Gross

Contamination can severely distort an estimator unless the estimation procedure is suitably robust. This is a well-known issue and has been addressed in Robust Statistics, however, the relation of contamination and distorted variable…

统计理论 · 数学 2022-07-15 Tino Werner

Biological sequencing data consist of read counts, e.g. of specified taxa and often exhibit sparsity (zero-count inflation) and overdispersion (extra-Poisson variability). As most sequencing techniques provide an arbitrary total count,…

应用统计 · 统计学 2024-07-01 Noora Kartiosuo , Jaakko Nevalainen , Olli Raitakari , Katja Pahkala , Kari Auranen

We examine the linear regression problem in a challenging high-dimensional setting with correlated predictors where the vector of coefficients can vary from sparse to dense. In this setting, we propose a combination of probabilistic…

统计方法学 · 统计学 2025-05-13 Roman Parzer , Peter Filzmoser , Laura Vana-Gür

Motivated by the challenges in analyzing gut microbiome and metagenomic data, this work aims to tackle the issue of measurement errors in high-dimensional regression models that involve compositional covariates. This paper marks a…

统计方法学 · 统计学 2024-09-13 Huali Zhao , Tianying Wang

By creating networks of biochemical pathways, communities of micro-organisms are able to modulate the properties of their environment and even the metabolic processes within their hosts. Next-generation high-throughput sequencing has led to…

应用统计 · 统计学 2023-03-28 Molly G. Hayes , Morgan G. I. Langille , Hong Gu

Identifying which taxa in our microbiota are associated with traits of interest is important for advancing science and health. However, the identification is challenging because the measured vector of taxa counts (by amplicon sequencing) is…

基因组学 · 定量生物学 2020-03-31 Barak Brill , Amnon Amir , Ruth Heller

Identifying relevant factors that influence the multinomial counts in compositional data is difficult in high dimensional settings due to the complex associations and overdispersion. Multivariate count models such as the…

统计方法学 · 统计学 2025-04-08 Alysha Cooper , Zeny Feng , Ayesha Ali , Tim Arciszewski , Lorna Deeth

Food authenticity studies are concerned with determining if food samples have been correctly labeled or not. Discriminant analysis methods are an integral part of the methodology for food authentication. Motivated by food authenticity…

统计方法学 · 统计学 2010-10-08 Thomas Brendan Murphy , Nema Dean , Adrian E. Raftery

Metagenomics sequencing is routinely applied to quantify bacterial abundances in microbiome studies, where the bacterial composition is estimated based on the sequencing read counts. Due to limited sequencing depth and DNA dropouts, many…

统计方法学 · 统计学 2019-04-26 Yuanpei Cao , Anru Zhang , Hongzhe Li

Quantile regression is useful for characterizing the conditional distribution of a response variable and understanding heterogeneity in the covariate effects at different quantiles. The rise of high-dimensional physiological data in…

统计方法学 · 统计学 2026-03-25 Yuanzhen Yue , Stella Self , Yichao Wu , Jiajia Zhang , Rahul Ghosal

Learning governing equations from a family of data sets which share the same physical laws but differ in bifurcation parameters is challenging. This is due, in part, to the wide range of phenomena that could be represented in the data sets…

数值分析 · 数学 2017-09-07 Hayden Schaeffer , Giang Tran , Rachel Ward

Dynamic treatment regimes (DTRs) consist of a sequence of decision rules, one per stage of intervention, that finds effective treatments for individual patients according to patient information history. DTRs can be estimated from models…

统计方法学 · 统计学 2021-12-07 Zeyu Bian , Erica EM Moodie , Susan M Shortreed , Sahir Bhatnagar

Diffusion-based generative models (DBGMs) perturb data to a target noise distribution and reverse this process to generate samples. The choice of noising process, or inference diffusion process, affects both likelihoods and sample quality.…

机器学习 · 计算机科学 2023-03-06 Raghav Singhal , Mark Goldstein , Rajesh Ranganath

The model interpretation is essential in many application scenarios and to build a classification model with a ease of model interpretation may provide useful information for further studies and improvement. It is common to encounter with a…

机器学习 · 统计学 2019-01-07 Wan-Ping Nicole Chen , Yuan-chin Ivan Chang

The Regression Discontinuity Design (RDD) is a quasi-experimental design that estimates the causal effect of a treatment when its assignment is defined by a threshold value for a continuous assignment variable. The RDD assumes that subjects…

应用统计 · 统计学 2020-03-27 Federico Ricciardi , Silvia Liverani , Gianluca Baio