中文
相关论文

相关论文: High heritability does not imply accurate predicti…

200 篇论文

Random forests are a statistical learning technique that use bootstrap aggregation to average high-variance and low-bias trees. Improvements to random forests, such as applying Lasso regression to the tree predictions, have been proposed in…

机器学习 · 统计学 2025-11-13 Jing Shang , James Bannon , Benjamin Haibe-Kains , Robert Tibshirani

Genetic variants (GVs) are defined as differences in the DNA sequences among individuals and play a crucial role in diagnosing and treating genetic diseases. The rapid decrease in next generation sequencing cost has led to an exponential…

机器学习 · 计算机科学 2024-12-06 Zehui Li , Vallijah Subasri , Guy-Bart Stan , Yiren Zhao , Bo Wang

To improve accuracy and speed of regressions and classifications, we present a data-based prediction method, Random Bits Regression (RBR). This method first generates a large number of random binary intermediate/derived features based on…

机器学习 · 统计学 2016-11-04 Yi Wang , Yi Li , Momiao Xiong , Li Jin

Gradient boosting of regression trees is a competitive procedure for learning predictive models of continuous data that fits the data with an additive non-parametric model. The classic version of gradient boosting assumes that the data is…

机器学习 · 计算机科学 2016-07-04 Iman Alodah , Jennifer Neville

The standard regression tree method applied to observations within clusters poses both methodological and implementation challenges. Effectively leveraging these data requires methods that account for both individual-level and sample-level…

统计方法学 · 统计学 2025-03-05 Jeremiah Allis , Xin Jin , Riddhi Ghosh

This paper characterizes the conditional distribution properties of the finite sample ridge regression estimator and uses that result to evaluate total regression and generalization errors that incorporate the inaccuracies committed at the…

机器学习 · 统计学 2016-05-31 Lyudmila Grigoryeva , Juan-Pablo Ortega

Given genetic variations and various phenotypical traits, such as Magnetic Resonance Imaging (MRI) features, we consider two important and related tasks in biomedical research: i)to select genetic and phenotypical markers for disease…

机器学习 · 计算机科学 2013-10-17 Shandian Zhe , Zenglin Xu , Yuan Qi

Symbolic regression is a nonlinear regression method which is commonly performed by an evolutionary computation method such as genetic programming. Quantification of uncertainty of regression models is important for the interpretation of…

机器学习 · 计算机科学 2022-09-15 Fabricio Olivetti de Franca , Gabriel Kronberger

Our work was motivated by a recent study on birth defects of infants born to pregnant women exposed to a certain medication for treating chronic diseases. Outcomes such as birth defects are rare events in the general population, which often…

应用统计 · 统计学 2017-02-24 Ronghui Xu , Jue Hou , Christina D. Chambers

The study of genomic variation has provided key insights into the functional role of mutations. Predominantly, studies have focused on single nucleotide variants (SNV), which are relatively easy to detect and can be described with rich…

基因组学 · 定量生物学 2015-09-04 Daniel R. Zerbino , Tracy Ballinger , Benedict Paten , Glenn Hickey , David Haussler

Exploring the genetic basis of heritable traits remains one of the central challenges in biomedical research. In simple cases, single polymorphic loci explain a significant fraction of the phenotype variability. However, many traits of…

种群与进化 · 定量生物学 2015-03-20 Barbara Rakitsch , Christoph Lippert , Oliver Stegle , Karsten Borgwardt

Understanding generalization and estimation error of estimators for simple models such as linear and generalized linear models has attracted a lot of attention recently. This is in part due to an interesting observation made in machine…

机器学习 · 统计学 2021-03-09 Mojtaba Sahraee-Ardakan , Tung Mai , Anup Rao , Ryan Rossi , Sundeep Rangan , Alyson K. Fletcher

For many traits, including susceptibility to common diseases in humans, causal loci uncovered by genetic mapping studies explain only a minority of the heritable contribution to trait variation. Multiple explanations for this "missing…

基因组学 · 定量生物学 2015-06-11 Joshua S. Bloom , Ian M. Ehrenreich , Wesley Loo , Thúy-Lan Võ Lite , Leonid Kruglyak

Traditional GWAS has advanced our understanding of complex diseases but often misses nonlinear genetic interactions. Deep learning offers new opportunities to capture complex genomic patterns, yet existing methods mostly depend on feature…

机器学习 · 计算机科学 2025-07-08 Iqra Farooq , Sara Atito , Ayse Demirkan , Inga Prokopenko , Muhammad Rana

The genetic etiologies of common diseases are highly complex and heterogeneous. Classic statistical methods, such as linear regression, have successfully identified numerous genetic variants associated with complex diseases. Nonetheless,…

应用统计 · 统计学 2020-10-28 Jinghang Lin , Xiaoran Tong , Chenxi Li , Qing Lu

A new method to improve the performance of Random weight change (RWC) algorithm based on a simple genetic algorithm, namely, Genetic random weight change (GRWC) is proposed. It is to find the optimal values of global minima via learning. In…

神经与进化计算 · 计算机科学 2019-06-06 Mohammad Ibraim Sarker , Yali Nie , Hong Yongki , Hyongsuk Kim

Linkage disequilibrium score regression (LDSC) has emerged as an essential tool for genetic and genomic analyses of complex traits, utilizing high-dimensional data derived from genome-wide association studies (GWAS). LDSC computes the…

统计方法学 · 统计学 2025-04-16 Fei Xue , Bingxin Zhao

Ridge regression is a well established regression estimator which can conveniently be adapted for classification problems. One compelling reason is probably the fact that ridge regression emits a closed-form solution thereby facilitating…

机器学习 · 计算机科学 2020-03-26 Jakramate Bootkrajang

Commonly, machine learning models minimize an empirical expectation. As a result, the trained models typically perform well for the majority of the data but the performance may deteriorate in less dense regions of the dataset. This issue…

机器学习 · 计算机科学 2021-07-22 Joachim Schreurs , Hannes De Meulemeester , Michaël Fanuel , Bart De Moor , Johan A. K. Suykens

We characterize and study variable importance (VIMP) and pairwise variable associations in binary regression trees. A key component involves the node mean squared error for a quantity we refer to as a maximal subtree. The theory naturally…

机器学习 · 统计学 2009-09-29 Hemant Ishwaran
‹ 上一页 1 8 9 10 下一页 ›