English
Related papers

Related papers: Enhancing Computational Efficiency in High-Dimensi…

200 papers

Heterogeneity is a fundamental characteristic of cancer. To accommodate heterogeneity, subgroup identification has been extensively studied and broadly categorized into unsupervised and supervised analysis. Compared to unsupervised…

Methodology · Statistics 2026-02-25 Xing Qin , Xu Liu , Shuangge Ma , Mengyun Wu

Gibbs sampling is a widely popular Markov chain Monte Carlo algorithm that can be used to analyze intractable posterior distributions associated with Bayesian hierarchical models. There are two standard versions of the Gibbs sampler: The…

Statistics Theory · Mathematics 2020-01-01 Grant Backlund , James P. Hobert , Yeun Ji Jung , Kshitij Khare

We consider the joint inference of regression coefficients and the inverse covariance matrix for covariates in high-dimensional probit regression, where the predictors are both relevant to the binary response and functionally related to one…

Methodology · Statistics 2022-03-15 Xuan Cao , Kyoungjae Lee

The use of Gaussian processes (GPs) is supported by efficient sampling algorithms, a rich methodological literature, and strong theoretical grounding. However, due to their prohibitive computation and storage demands, the use of exact GPs…

Statistics Theory · Mathematics 2022-07-27 Kelly R. Moran , Matthew W. Wheeler

Applications of high-dimensional regression often involve multiple sources or types of covariates. We propose methodology for this setting, emphasizing the "wide data" regime with large total dimensionality p and sample size n<<p. We focus…

Bayesian feature allocation models are a popular tool for modelling data with a combinatorial latent structure. Exact inference in these models is generally intractable and so practitioners typically apply Markov Chain Monte Carlo (MCMC)…

Computation · Statistics 2020-01-28 Alexandre Bouchard-Côté , Andrew Roth

The main objective of this paper is to apply linear and pretest shrinkage estimation techniques to estimating the parameters of two 2-parameter Burr-XII distributions. Further more, predictions for future observations are made using both…

Methodology · Statistics 2024-01-09 Soheila Akbari Bargoshadi , Hossein Bevrani

Next-generation sequencing technologies provide a revolutionary tool for generating gene expression data. Starting with a fixed RNA sample, they construct a library of millions of differentially abundant short sequence tags or "reads",…

Quantitative Methods · Quantitative Biology 2014-05-13 Dimitrios V. Vavoulis , Julian Gough

Simultaneous analysis of gene expression data and genetic variants is highly of interest, especially when the number of gene expressions and genetic variants are both greater than the sample size. Association of both causal genes and…

Methodology · Statistics 2021-10-07 Morteza Amini

Bayesian optimization is an effective methodology for the global optimization of functions with expensive evaluations. It relies on querying a distribution over functions defined by a relatively cheap surrogate model. An accurate model for…

Bayesian variable selection is a powerful tool for data analysis, as it offers a principled method for variable selection that accounts for prior information and uncertainty. However, wider adoption of Bayesian variable selection has been…

Methodology · Statistics 2022-09-13 Martin Jankowiak

Identifying genes underlying cancer development is critical to cancer biology and has important implications across prevention, diagnosis and treatment. Cancer sequencing studies aim at discovering genes with high frequencies of somatic…

Applications · Statistics 2013-12-09 Jie Ding , Lorenzo Trippa , Xiaogang Zhong , Giovanni Parmigiani

In this vignette, we introduce the UPG package for efficient Bayesian inference in probit, logit, multinomial logit and binomial logit models. UPG offers a convenient estimation framework for balanced and imbalanced data settings where…

Computation · Statistics 2023-07-03 Gregor Zens , Sylvia Frühwirth-Schnatter , Helga Wagner

Estimating the sharing of genetic effects across different conditions is important to many statistical analyses of genomic data. The patterns of sharing arising from these data are often highly heterogeneous. To flexibly model these…

Methodology · Statistics 2024-06-14 Yunqi Yang , Peter Carbonetto , David Gerard , Matthew Stephens

Breast cancer is the most common cancer among women worldwide. Early-stage diagnosis of breast cancer can significantly improve the efficiency of treatment. Computer-aided diagnosis (CAD) systems are widely adopted in this issue due to…

Image and Video Processing · Electrical Eng. & Systems 2024-10-28 Mohammad Reza Abbasniya , Sayed Ali Sheikholeslamzadeh , Hamid Nasiri , Samaneh Emami

The problem of joint estimation of multiple graphical models from high dimensional data has been studied in the statistics and machine learning literature, due to its importance in diverse fields including molecular biology, neuroscience…

Methodology · Statistics 2019-07-04 Peyman Jalali , Kshitij Khare , George Michailidis

Bi-clustering is a useful approach in analyzing biological data when observations come from heterogeneous groups and have a large number of features. We outline a general Bayesian approach in tackling bi-clustering problems in moderate to…

Applications · Statistics 2021-02-11 Han Yan , Jiexing Wu , Yang Li , Jun S. Liu

Motivated by the increasing use of and rapid changes in array technologies, we consider the prediction problem of fitting a linear regression relating a continuous outcome $Y$ to a large number of covariates $\mathbf {X}$, for example,…

Applications · Statistics 2014-01-13 Philip S. Boonstra , Bhramar Mukherjee , Jeremy M. G. Taylor

With the advancements of computer architectures, the use of computational models proliferates to solve complex problems in many scientific applications such as nuclear physics and climate research. However, the potential of such models is…

Computation · Statistics 2021-07-05 Vojtech Kejzlar , Tapabrata Maiti

Bayesian sample size calculations in clinical trials usually rely on complex Monte Carlo simulations in practice. Obtaining bounds on Bayesian notions of the false-positive rate and power often lack closed-form or approximate numerical…

Methodology · Statistics 2026-03-03 Riko Kelter