中文
相关论文

相关论文: variable selection and missing data imputation in …

200 篇论文

Bayesian variable selection is a powerful tool for data analysis, as it offers a principled method for variable selection that accounts for prior information and uncertainty. However, wider adoption of Bayesian variable selection has been…

统计方法学 · 统计学 2022-09-13 Martin Jankowiak

Random forest (RF) missing data algorithms are an attractive approach for dealing with missing data. They have the desirable properties of being able to handle mixed types of missing data, they are adaptive to interactions and nonlinearity,…

机器学习 · 统计学 2017-01-23 Fei Tang , Hemant Ishwaran

Large-scale genome-wide association studies (GWAS) have offered an exciting opportunity to discover putative causal genes or risk factors associated with diseases by using SNPs as instrumental variables (IVs). However, conventional…

统计方法学 · 统计学 2023-10-27 Ben Dai , Chunlin Li , Haoran Xue , Wei Pan , Xiaotong Shen

In longitudinal data a response variable is measured over time, or under different conditions, for a cohort of individuals. In many situations all intended measurements are not available which results in missing values. If the missing value…

统计方法学 · 统计学 2022-08-10 Ahmed M. Gad , Nesma M. Darwish

Genome-wide association studies generate very large datasets that require scalable analysis algorithms. In this report we describe the GEDI software package, which implements efficient algorithms for performing several common tasks in the…

数据结构与算法 · 计算机科学 2016-09-08 Justin Kennedy , Ion I. Mandoiu , Bogdan Pasaniuc

Cancer prognosis is often based on a set of omics covariates and a set of established clinical covariates such as age and tumor stage. Combining these two sets poses challenges. First, dimension difference: clinical covariates should be…

统计方法学 · 统计学 2024-11-05 Jeroen M. Goedhart , Mark A. van de Wiel , Wessel N. van Wieringen , Thomas Klausch

The broad sense genetic heritability, which quantifies the total proportion of phenotypic variation in a population due to genetic factors, is crucial for understanding trait inheritance. While many existing methods focus on estimating…

统计方法学 · 统计学 2024-11-04 Olivia Bley , Elizabeth Lei , Andy Zhou , Xiaoxi Shen

This paper addresses the challenge of efficiently capturing a high proportion of true signals for subsequent data analyses when sample sizes are relatively limited with respect to data dimension. We propose the signal missing rate as a new…

统计方法学 · 统计学 2018-08-30 X. Jessie Jeng , Teng Zhang , Jung-Ying Tzeng

We propose a new method to impute missing values in mixed datasets. It is based on a principal components method, the factorial analysis for mixed data, which balances the influence of all the variables that are continuous and categorical…

应用统计 · 统计学 2013-02-20 Vincent Audigier , François Husson , Julie Josse

We present an alternative method for genome-wide association studies (GWAS) that is more powerful than the regular GWAS method for locus detection. The regular GWAS method suffers from a substantial multiple-testing burden because of the…

It is generally acknowledged that most complex diseases are affected in part by interactions between genes and genes and/or between genes and environmental factors. Taking into account environmental exposures and their interactions with…

应用统计 · 统计学 2014-06-19 Flora Alarcon , Vittorio Perduca , Gregory Nuel

Research on the localization of the genetic basis associated with diseases or traits has been widely conducted in the last a few decades. Scan methods have been developed for region-based analysis in whole-genome association studies,…

统计方法学 · 统计学 2024-10-31 Wei Zhang , Fan Wang , Fang Yao

Global expression analyses using microarray technologies are becoming more common in genomic research, therefore, new statistical challenges associated with combining information from multiple studies must be addressed. In this paper we…

应用统计 · 统计学 2013-01-29 Jia Li , George C. Tseng

Motivated by the CATHGEN data, we develop a new statistical learning method for simultaneous variable selection and parameter estimation under the context of generalized partly linear models for data with high-dimensional covariates. The…

统计方法学 · 统计学 2023-11-02 Christian Chan , Xiaotian Dai , Thierry Chekouo , Quan Long , Xuewen Lu

A large number of recent genome-wide association studies (GWASs) for complex phenotypes confirm the early conjecture for polygenicity, suggesting the presence of large number of variants with only tiny or moderate effects. However, due to…

基因组学 · 定量生物学 2018-05-01 Mingwei Dai , Xiang Wan , Hao Peng , Yao Wang , Yue Liu , Jin Liu , Zongben Xu , Can Yang

The variance component tests used in genomewide association studies of thousands of individuals become computationally exhaustive when multiple traits are analysed in the context of omics studies. We introduce two high-throughput algorithms…

计算工程、金融与科学 · 计算机科学 2012-11-13 Diego Fabregat-Traver , Yurii S. Aulchenko , Paolo Bientinesi

Drug development is a very costly and lengthy process, while repositioned or repurposed drugs could be brought into clinical practice within a shorter time-frame and at a much reduced cost. The past decade has observed a massive growth in…

基因组学 · 定量生物学 2019-11-14 Alexandria Lau , Hon-Cheong So

We develop a general statistical framework for the analysis and inference of large tree-structured data, with a focus on developing asymptotic goodness-of-fit tests. We first propose a consistent statistical model for binary trees, from…

We propose to learn latent graphical models when data have mixed variables and missing values. This model could be used for further data analysis, including regression, classification, ranking etc. It also could be used for imputing missing…

统计方法学 · 统计学 2015-11-17 Xiao Li , Jinzhu Jia , Yuan Yao

Motivated by applications in neuroanatomy, we propose a novel methodology for estimating the heritability which corresponds to the proportion of phenotypic variance which can be explained by genetic factors. Estimating this quantity for…