English
Related papers

Related papers: variable selection and missing data imputation in …

200 papers

In this paper we have proposed a model for the distribution of allelic probabilities for generating populations as reliably as possible. Our objective was to develop such a model which would allow simulating allelic probabilities with…

Association testing aims to discover the underlying relationship between genotypes (usually Single Nucleotide Polymorphisms, or SNPs) and phenotypes (attributes, or traits). The typically large data sets used in association testing often…

Applications · Statistics 2012-07-04 Zhen Li , Vikneswaran Gopal , Xiaobo Li , John M. Davis , George Casella

Linkage disequilibrium score regression (LDSC) has emerged as an essential tool for genetic and genomic analyses of complex traits, utilizing high-dimensional data derived from genome-wide association studies (GWAS). LDSC computes the…

Methodology · Statistics 2025-04-16 Fei Xue , Bingxin Zhao

In genome-wide association studies (GWAS), penalization is an important approach for identifying genetic markers associated with trait while mixed model is successful in accounting for a complicated dependence structure among samples.…

Methodology · Statistics 2013-05-21 Jin Liu , Can Yang , Xingjie Shi , Cong Li , Jian Huang , Hongyu Zhao , Shuangge Ma

Transcriptome-wide association studies (TWAS) are powerful tools for identifying gene-level associations by integrating genome-wide association studies and gene expression data. However, most TWAS methods focus on linear associations…

Methodology · Statistics 2024-12-10 Tianying Wang , Iuliana Ionita-Laza , Ying Wei

We introduce a statistical method that can reconstruct nonlinear genetic models (i.e., including epistasis, or gene-gene interactions) from phenotype-genotype (GWAS) data. The computational and data resource requirements are similar to…

Genomics · Quantitative Biology 2015-09-29 Chiu Man Ho , Stephen D. H. Hsu

The central aim in this paper is to address variable selection questions in nonlinear and nonparametric regression. Motivated by statistical genetics, where nonlinear interactions are of particular interest, we introduce a novel and…

Methodology · Statistics 2018-08-28 Lorin Crawford , Seth R. Flaxman , Daniel E. Runcie , Mike West

Traditional GWAS has advanced our understanding of complex diseases but often misses nonlinear genetic interactions. Deep learning offers new opportunities to capture complex genomic patterns, yet existing methods mostly depend on feature…

Machine Learning · Computer Science 2025-07-08 Iqra Farooq , Sara Atito , Ayse Demirkan , Inga Prokopenko , Muhammad Rana

Handling missing data is a major challenge in model-based clustering, especially when the data exhibit skewness and heavy tails. We address this by extending the finite mixture of scale mixtures of multivariate skew-normal (FMSMSN) family…

Methodology · Statistics 2025-07-29 Jason Pillay , Cristina Tortora , Antonio Punzo , Andriette Bekker

Blockwise missing data occurs frequently when we integrate multisource or multimodality data where different sources or modalities contain complementary information. In this paper, we consider a high-dimensional linear regression model with…

Methodology · Statistics 2023-06-30 Fei Xue , Rong Ma , Hongzhe Li

Random Forests [Breiman:2001] (RF) are a fully non-parametric statistical method requiring no distributional assumptions on covariate relation to the response. RF are a robust, nonlinear technique that optimizes predictive accuracy by…

Computation · Statistics 2016-12-30 John Ehrlinger

Standard approaches to tackle high-dimensional supervised classification problem often include variable selection and dimension reduction procedures. The novel methodology proposed in this paper combines clustering of variables and feature…

Statistics Theory · Mathematics 2018-11-07 Marie Chavent , Robin Genuer , Jerome Saracco

Data from discovery proteomic and phosphoproteomic experiments typically include missing values that correspond to proteins that have not been identified in the analyzed sample. Replacing the missing values with random numbers, a process…

Quantitative Methods · Quantitative Biology 2019-10-01 Matus Medo , Daniel M. Aebersold , Michaela Medova

One of the notable fields in studying the genetics of cancer is disease gene identification which affects disease treatment and drug discovery. Many researches have been done in this field. Genome-wide association studies (GWAS) are one of…

Computational Engineering, Finance, and Science · Computer Science 2016-04-27 Zahra Razaghi-Moghadama , Razieh Abdollahia , Sama Goliaeib , Morteza Ebrahimia

In this paper, a robust weighted score for unbalanced data (ROWSU) is proposed for selecting the most discriminative feature for high dimensional gene expression binary classification with class-imbalance problem. The method addresses one…

Machine Learning · Statistics 2024-01-24 Zardad Khan , Amjad Ali , Saeed Aldahmani

Combining data from several case-control genome-wide association (GWA) studies can yield greater efficiency for detecting associations of disease with single nucleotide polymorphisms (SNPs) than separate analyses of the component studies.…

Methodology · Statistics 2010-10-26 Ruth M. Pfeiffer , Mitchell H. Gail , David Pee

Following the publication of an attack on genome-wide association studies (GWAS) data proposed by Homer et al., considerable attention has been given to developing methods for releasing GWAS data in a privacy-preserving way. Here, we…

Machine Learning · Statistics 2014-07-31 Fei Yu , Michal Rybar , Caroline Uhler , Stephen E. Fienberg

Exploratory data analysis is crucial for developing and understanding classification models from high-dimensional datasets. We explore the utility of a new unsupervised tree ensemble called uncharted forest for visualizing class…

Machine Learning · Statistics 2018-07-03 Casey Kneale , Steven D. Brown

Big Data is one of the major challenges of statistical science and has numerous consequences from algorithmic and theoretical viewpoints. Big Data always involve massive data but they also often include online data and data heterogeneity.…

Machine Learning · Statistics 2017-03-23 Robin Genuer , Jean-Michel Poggi , Christine Tuleau-Malot , Nathalie Villa-Vialaneix

Tree-based ensemble methods, as Random Forests and Gradient Boosted Trees, have been successfully used for regression in many applications and research studies. Furthermore, these methods have been extended in order to deal with uncertainty…

Machine Learning · Computer Science 2018-11-20 Myriam Tami , Marianne Clausel , Emilie Devijver , Adrien Dulac , Eric Gaussier , Stefan Janaqi , Meriam Chebre