English
Related papers

Related papers: Algorithms for Large-scale Whole Genome Associatio…

200 papers

The recent super-exponential growth in the amount of sequencing data generated worldwide has put techniques for compressed storage into the focus. Most available solutions, however, are strictly tied to specific bioinformatics formats,…

Genomics · Quantitative Biology 2021-11-01 Łukasz Roguski , Paolo Ribeca

There is little debate about the importance of the ancestral recombination graph in population genetics. An important theoretical tool, the main obstacle to its widespread usage is the computational cost required to match the…

Populations and Evolution · Quantitative Biology 2026-05-14 Patrick Fournier , Fabrice Larribe

We present a coherent Bayesian framework for selection of the most likely model from the five genetic models (genotypic, additive, dominant, co-dominant, and recessive) commonly used in genetic association studies. The approach uses a…

Methodology · Statistics 2015-04-22 Harold Bae , Thomas Perls , Martin Steinberg , Paola Sebastiani

The variance component tests used in genomewide association studies of thousands of individuals become computationally exhaustive when multiple traits are analysed in the context of omics studies. We introduce two high-throughput algorithms…

Computational Engineering, Finance, and Science · Computer Science 2012-11-13 Diego Fabregat-Traver , Yurii S. Aulchenko , Paolo Bientinesi

Large spatial datasets are becoming ubiquitous in environmental sciences with the explosion in the amount of data produced by sensors that monitor and measure the Earth system. Consequently, the geostatistical analysis of these data…

Statistics Theory · Mathematics 2018-06-06 Thomas Romary , Nicolas Desassis

Graphs and networks are common ways of depicting biological information. In biology, many different biological processes are represented by graphs, such as regulatory networks, metabolic pathways and protein--protein interaction networks.…

Applications · Statistics 2010-11-16 Caiyan Li , Hongzhe Li

Modelling gene-gene epistatic interactions when computing genetic risk scores is not a well-explored subfield of genetics and could have potential to improve risk stratification in practice. Though applications of machine learning (ML) show…

Genomics · Quantitative Biology 2023-06-16 Nathaniel Gunter , Prashanthi Vemuri , Vijay Ramanan , Robel K Gebre

Estimation and hypothesis tests for the covariance matrix in high dimensions is a challenging problem as the traditional multivariate asymptotic theory is no longer valid. When the dimension is larger than or increasing with the sample…

Methodology · Statistics 2020-11-18 Deepak Nag Ayyala , Santu Ghosh , Daniel F. Linder

It is common to conduct causal inference in matched observational studies by proceeding as though treatment assignments within matched sets are assigned uniformly at random and using this distribution as the basis for inference. This…

Methodology · Statistics 2023-11-14 Samuel D. Pimentel , Yaxuan Huang

Recently, Deep Neural Networks (DNNs) have recorded great success in handling medical and other complex classification tasks. However, as the sizes of a DNN model and the available dataset increase, the training process becomes more complex…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-02-08 Samson B. Akintoye , Liangxiu Han , Xin Zhang , Haoming Chen , Daoqiang Zhang

1. Deciphering coexistence patterns is a current challenge to understanding diversity maintenance, especially in rich communities where the complexity of these patterns is magnified through indirect interactions that prevent their…

Machine Learning · Computer Science 2021-07-14 J. Hirn , J. E. García , A. Montesinos-Navarro , R. Sanchez-Martín , V. Sanz , M. Verdú

One of the major challenges in coreference resolution is how to make use of entity-level features defined over clusters of mentions rather than mention pairs. However, coreferent mentions usually spread far apart in an entire text, which…

Computation and Language · Computer Science 2023-07-25 Lu Liu , Zhenqiao Song , Xiaoqing Zheng , Jun He

The widely used genetic pleiotropic analysis of multiple phenotypes are often designed for examining the relationship between common variants and a few phenotypes. They are not suited for both high dimensional phenotypes and high…

Machine Learning · Statistics 2015-12-04 Panpan Wang , Mohammad Rahman , Li Jin , Momiao Xiong

Combined inference for heterogeneous high-dimensional data is critical in modern biology, where clinical and various kinds of molecular data may be available from a single study. Classical genetic association studies regress a single…

Applications · Statistics 2017-03-22 Hélène Ruffieux , Anthony C. Davison , Jörg Hager , Irina Irincheeva

In systems biomedicine, an experimenter encounters different potential sources of variation in data such as individual samples, multiple experimental conditions, and multi-variable network-level responses. In multiparametric cytometry,…

The problem of testing changes in covariance has received increasing attention in recent years, especially in the context of high-dimensional testing. A number of approaches have been proposed, all limited to the two-sample problem and…

Methodology · Statistics 2016-09-06 Yi-Hui Zhou

Machine Learning methods have of late made significant efforts to solving multidisciplinary problems in the field of cancer classification using microarray gene expression data. Feature subset selection methods can play an important role in…

Computational Engineering, Finance, and Science · Computer Science 2013-03-04 G. Prat , Ll. Belanche

Gene-gene interactions play a crucial role in the manifestation of complex human diseases. Uncovering significant gene-gene interactions is a challenging task. Here, we present an innovative approach utilizing data-driven computational…

Artificial Intelligence · Computer Science 2024-10-22 Yifan Wu , Yuntao Yang , Zirui Liu , Zhao Li , Khushbu Pahwa , Rongbin Li , Wenjin Zheng , Xia Hu , Zhaozhuo Xu

Polygenic risk scores and other genomic analyses require large individual-level genotype datasets, yet strict data access restrictions impede sharing. Synthetic genotype generation offers a privacy-preserving alternative, but most existing…

We provide a view on high-dimensional statistical inference for genome-wide association studies (GWAS). It is in part a review but covers also new developments for meta analysis with multiple studies and novel software in terms of an…

Applications · Statistics 2020-02-17 Claude Renaux , Laura Buzdugan , Markus Kalisch , Peter Bühlmann