中文
相关论文

相关论文: Variable selection for sparse Dirichlet-multinomia…

200 篇论文

We consider a novel Bayesian approach to estimation, uncertainty quantification, and variable selection for a high-dimensional linear regression model under sparsity. The number of predictors can be nearly exponentially large relative to…

统计方法学 · 统计学 2025-02-28 Samhita Pal , Subhashis Ghoshal

In this paper, we propose Varying Effects Regression with Graph Estimation (VERGE), a novel Bayesian method for feature selection in regression. Our model has key aspects that allow it to leverage the complex structure of data sets arising…

统计方法学 · 统计学 2024-10-10 Yangfan Ren , Christine B. Peterson , Marina Vannucci

In multivariate regression, when covariates are numerous, it is often reasonable to assume that only a small number of them has predictive information. In some medical applications for instance, it is believed that only a few genes out of…

统计方法学 · 统计学 2022-07-12 Sylvain Sardy , Xiaoyu Ma

Many biological high-throughput data sets, such as targeted amplicon-based and metagenomic sequencing data, are compositional in nature. A common exploratory data analysis task is to infer statistical associations between the…

统计方法学 · 统计学 2020-07-28 Aditya Mishra , Christian L. Muller

In microbiome studies, it is of interest to use a sample from a population of microbes, such as the gut microbiota community, to estimate the population proportion of these taxa. However, due to biases introduced in sampling and…

统计方法学 · 统计学 2022-10-11 Roulan Jiang , Xiang Zhan , Tianying Wang

In high-dimensional statistics, variable selection recovers the latent sparse patterns from all possible covariate combinations. This paper proposes a novel optimization method to solve the exact L0-regularized regression problem, which is…

统计方法学 · 统计学 2022-06-02 Mingzhang Yin , Nhat Ho , Bowei Yan , Xiaoning Qian , Mingyuan Zhou

This paper studies the case of possibly high-dimensional covariates in the regression discontinuity design (RDD) analysis. In particular, we propose estimation and inference methods for the RDD models with covariate selection which perform…

计量经济学 · 经济学 2026-01-21 Yoichi Arai , Taisuke Otsu , Myung Hwan Seo

Motivated by an application in high-throughput genomics and metabolomics, we propose a novel, efficient and fully data-driven approach for estimating large block structured sparse covariance matrices in the case where the number of…

统计方法学 · 统计学 2019-12-09 Marie Perrot-Dockès , Céline Lévy-Leduc , Loïc Rajjou

The interpretation of count data originating from the current generation of DNA sequencing platforms requires special attention. In particular, the per-sample library sizes often vary by orders of magnitude from the same sequencing run, and…

定量方法 · 定量生物学 2015-06-17 Paul J. McMurdie , Susan Holmes

Distributed statistical learning has become a popular technique for large-scale data analysis. Most existing work in this area focuses on dividing the observations, but we propose a new algorithm, DDAC-SpAM, which divides the features under…

机器学习 · 计算机科学 2023-07-11 Yifan He , Ruiyang Wu , Yong Zhou , Yang Feng

We present a new variable selection method based on model-based gradient boosting and randomly permuted variables. Model-based boosting is a tool to fit a statistical model while performing variable selection at the same time. A drawback of…

机器学习 · 统计学 2017-02-16 Janek Thomas , Tobias Hepp , Andreas Mayr , Bernd Bischl

Inferring concerted changes among biological traits along an evolutionary history remains an important yet challenging problem. Besides adjusting for spurious correlation induced from the shared history, the task also requires sufficient…

Nonlinear Mixed effects models are hidden variables models that are widely used in many fields such as pharmacometrics. In such models, the distribution characteristics of hidden variables can be specified by including several parameters…

统计方法学 · 统计学 2021-10-19 Edouard Ollier

Variable selection naturally arises as a useful subject when faced with data with massive predictor space. In addition to the massive dimensionality, the data may be characterized by intra-subject correlation, and cure fraction, which are…

统计方法学 · 统计学 2025-12-24 Richard Tawiah , Shu Kay Ng , Geoffrey J. McLachlan

Microbial communities analysis is drawing growing attention due to the rapid development of high-throughput sequencing techniques nowadays. The observed data has the following typical characteristics: it is high-dimensional, compositional…

统计方法学 · 统计学 2020-04-30 Yong He , Pengfei Liu , Xinsheng Zhang , Wang Zhou

Many studies have been performed to characterize the dynamics and stability of the microbiome across a range of environmental contexts [Costello et al., 2012, Faust et al., 2015]. For example, it is often of interest to identify time…

应用统计 · 统计学 2017-12-04 Kris Sankaran , Susan P. Holmes

Many scientific and industrial processes produce data that is best analysed as vectors of relative values, often called compositions or proportions. The Dirichlet distribution is a natural distribution to use for composition or proportion…

统计方法学 · 统计学 2020-04-15 Sean van der Merwe

A method is introduced to perform simultaneous sparse dimension reduction on two blocks of variables. Beyond dimension reduction, it also yields an estimator for multivariate regression with the capability to intrinsically deselect…

统计方法学 · 统计学 2024-11-28 Sven Serneels

Communication and coordination play a major role in the ability of bacterial cells to adapt to ever changing environments and conditions. Recent work has shown that such coordination underlies several aspects of bacterial responses…

The gut microbiome plays a crucial role in human health, yet the mechanisms underlying host-microbiome interactions remain unclear, limiting its translational potential. Recent microbiome multiomics studies, particularly paired…

统计方法学 · 统计学 2025-04-09 Haoran Shi , Yue Wang , Dan Cheng