中文
相关论文

相关论文: A robust statistical method for Genome-wide associ…

200 篇论文

Two-sample testing is a fundamental problem in statistics. Despite its long history, there has been renewed interest in this problem with the advent of high-dimensional and complex data. Specifically, in the machine learning literature,…

统计方法学 · 统计学 2019-11-19 Ilmun Kim , Ann B. Lee , Jing Lei

It is generally acknowledged that most complex diseases are affected in part by interactions between genes and genes and/or between genes and environmental factors. Taking into account environmental exposures and their interactions with…

应用统计 · 统计学 2014-06-19 Flora Alarcon , Vittorio Perduca , Gregory Nuel

Few Bayesian methods for analyzing high-dimensional sparse survival data provide scalable variable selection, effect estimation and uncertainty quantification. Such methods often either sacrifice uncertainty quantification by computing…

统计方法学 · 统计学 2022-07-06 Michael Komodromos , Eric Aboagye , Marina Evangelou , Sarah Filippi , Kolyan Ray

An innovative sampling strategy is proposed, which applies to large-scale population-based surveys targeting a rare trait that is unevenly spread over a geographical area of interest. Our proposal is characterised by the ability to tailor…

统计方法学 · 统计学 2020-04-07 Fulvia Mecatti , Charalambos Sismanidis , Emanuela Furfaro

Gene regulatory networks play a crucial role in controlling an organism's biological processes, which is why there is significant interest in developing computational methods that are able to extract their structure from high-throughput…

机器学习 · 统计学 2018-09-19 Ioan Gabriel Bucur , Tom van Bussel , Tom Claassen , Tom Heskes

Canonical correlation analysis (CCA) is a classic statistical method for discovering latent co-variation that underpins two or more observed random vectors. Several extensions and variations of CCA have been proposed that have strengthened…

机器学习 · 计算机科学 2023-12-22 Paris A. Karakasis , Nicholas D. Sidiropoulos

Cellwise outliers are widespread in data and traditional robust methods may fail when applied to datasets under such contamination. We propose a variable selection procedure, that uses a pairwise robust estimator to obtain an initial…

统计方法学 · 统计学 2023-09-06 Peng Su , Garth Tarr , Samuel Muller

Motivated by empirical arguments that are well-known from the genome-wide association studies (GWAS) literature, we study the statistical properties of linear mixed models (LMMs) applied to GWAS. First, we study the sensitivity of LMMs to…

定量方法 · 定量生物学 2021-11-09 Haohan Wang , Bryon Aragam , Eric Xing

Several statistical models used in genome-wide prediction assume independence of marker allele substitution effects, but it is known that these effects might be correlated. In statistics, graphical models have been identified as a useful…

定量方法 · 定量生物学 2017-04-13 Carlos Alberto Martínez , Kshitij Khare , Syed Rahman , Mauricio A. Elzo

Variational Bayes (VB) is rapidly becoming a popular tool for Bayesian inference in statistical modeling. However, the existing VB algorithms are restricted to cases where the likelihood is tractable, which precludes the use of VB in many…

统计方法学 · 统计学 2016-08-05 Minh-Ngoc Tran , David J. Nott , Robert Kohn

The Critical Assessment of Genome Interpretation (CAGI) aims to advance the state of the art for computational prediction of genetic variant impact, particularly those relevant to disease. The five complete editions of the CAGI community…

基因组学 · 定量生物学 2024-06-19 The Critical Assessment of Genome Interpretation Consortium

Substantial progress has been made in identifying single genetic variants predisposing to common complex diseases. Nonetheless, the genetic etiology of human diseases remains largely unknown. Human complex diseases are likely influenced by…

统计方法学 · 统计学 2014-05-27 Zihuai He , Min Zhang , Xiaowei Zhan , Qing Lu

In the context of survival analysis, data-driven neural network-based methods have been developed to model complex covariate effects. While these methods may provide better predictive performance than regression-based approaches, not all…

机器学习 · 统计学 2024-04-23 Jesse Islam , Maxime Turgeon , Robert Sladek , Sahir Bhatnagar

The challenge of Out-of-Distribution (OOD) generalization poses a foundational concern for the application of machine learning algorithms to risk-sensitive areas. Inspired by traditional importance weighting and propensity weighting…

机器学习 · 计算机科学 2025-02-12 Han Yu , Yue He , Renzhe Xu , Dongbai Li , Jiayin Zhang , Wenchao Zou , Peng Cui

In many practices, scientists are particularly interested in detecting which of the predictors are truly associated with a multivariate response. It is more accurate to model multiple responses as one vector rather than separating each…

统计方法学 · 统计学 2021-11-16 Xiaotian Dai , Guifang Fu , Randall Reese , Shaofei Zhao , Zuofeng Shang

We propose a resampling-based fast variable selection technique for detecting relevant single nucleotide polymorphisms (SNP) in a multi-marker mixed effect model. Due to computational complexity, current practice primarily involves testing…

应用统计 · 统计学 2025-04-30 Subhabrata Majumdar , Saonli Basu , Matt McGue , Snigdhansu Chatterjee

We propose a Bayesian variable selection method in the framework of modal regression for heavy-tailed responses. An efficient expectation-maximization algorithm is employed to expedite parameter estimation. A test statistic is constructed…

统计方法学 · 统计学 2025-10-29 Jiasong Duan , Hongmei Zhang , Xianzheng Huang

Global expression analyses using microarray technologies are becoming more common in genomic research, therefore, new statistical challenges associated with combining information from multiple studies must be addressed. In this paper we…

应用统计 · 统计学 2013-01-29 Jia Li , George C. Tseng

Summary statistics of genome-wide association studies (GWAS) teach causal relationship between millions of genetic markers and tens and thousands of phenotypes. However, underlying biological mechanisms are yet to be elucidated. We can…

机器学习 · 统计学 2019-01-25 Yongjin Park , Abhishek Sarkar , Khoi Nguyen , Manolis Kellis

Categorical variables are of uttermost importance in biomedical research. When two of them are considered, it is often the case that one wants to test whether or not they are statistically dependent. We show weaknesses of classical methods…