English
Related papers

Related papers: High Performance Solutions for Big-data GWAS

200 papers

Motivation: Genome-wide association studies (GWASs), which assay more than a million single nucleotide polymorphisms (SNPs) in thousands of individuals, have been widely used to identify genetic risk variants for complex diseases. However,…

Computational Engineering, Finance, and Science · Computer Science 2015-01-27 Ben Teng , Can Yang , Jiming Liu , Zhipeng Cai , Xiang Wan

Federated Learning has gained attention for its ability to enable multiple nodes to collaboratively train machine learning models without sharing raw data. At the same time, Generative AI -- particularly Generative Adversarial Networks…

Machine Learning · Computer Science 2026-01-19 Youssef Tawfilis , Hossam Amer , Minar El-Aasser , Tallal Elshabrawy

Genome-wide association studies (GWAS) suggests that a complex disease is typically affected by many genetic variants with small or moderate effects. Identification of these risk variants remains to be a very challenging problem.…

Methodology · Statistics 2014-01-21 Dongjun Chung , Can Yang , Cong Li , Joel Gelernter , Hongyu Zhao

In systems biology, it is common to measure biochemical entities at different levels of the same biological system. One of the central problems for the data fusion of such data sets is the heterogeneity of the data. This thesis discusses…

Genomics · Quantitative Biology 2019-08-27 Yipeng Song

Computing the matching statistics of patterns with respect to a text is a fundamental task in bioinformatics, but a formidable one when the text is a highly compressed genomic database. Bannai et al. gave an efficient solution for this…

Data Structures and Algorithms · Computer Science 2021-02-12 Christina Boucher , Travis Gagie , Tomohiro I , Dominik Köppl , Ben Langmead , Giovanni Manzini , Gonzalo Navarro , Alejandro Pacheco , Massimiliano Rossi

Capturing the vast amount of meaningful information encoded in the human genome is a fascinating research problem. The outcome of these researches have significant influences in a number of health related fields --- personalized medicine,…

Cryptography and Security · Computer Science 2017-03-07 Mohammad Zahidul Hasan , Md Safiur Rahman Mahdi , Noman Mohammed

The data association problem is concerned with separating data coming from different generating processes, for example when data come from different data sources, contain significant noise, or exhibit multimodality. We present a fully…

Machine Learning · Statistics 2019-05-07 Markus Kaiser , Clemens Otte , Thomas Runkler , Carl Henrik Ek

The objective of a genome-wide association study (GWAS) is to associate subsequences of individuals' genomes to the observable characteristics called phenotypes (e.g., high blood pressure). Motivated by the GWAS problem, in this paper we…

Information Theory · Computer Science 2020-10-15 Behrooz Tahmasebi , Mohammad Ali Maddah-Ali , Seyed Abolfazl Motahari

In order to fully utilize "big data", it is often required to use "big models". Such models tend to grow with the complexity and size of the training data, and do not make strong parametric assumptions upfront on the nature of the…

Machine Learning · Statistics 2015-04-17 Vikas Sindhwani , Haim Avron

The capability of classifying and clustering a desired set of data is an essential part of building knowledge from data. However, as the size and dimensionality of input data increases, the run-time for such clustering algorithms is…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-07-25 Hadi Mardani Kamali

Emerging applications of machine learning in numerous areas involve continuous gathering of and learning from streams of data. Real-time incorporation of streaming data into the learned models is essential for improved inference in these…

Machine Learning · Computer Science 2020-12-01 Matthew Nokleby , Haroon Raja , Waheed U. Bajwa

The problem addressed here is that of simultaneous treatment of several gene expression datasets, possibly collected under different experimental conditions and/or platforms. Using robust statistics, a large scale statistical analysis has…

Methodology · Statistics 2014-10-10 Bernard Ycart , Konstantina Charmpi , Sophie Rousseaux , Jean-Jacques Fournié

Motivation: Modern bioinformatics workflows, particularly in imaging and representation learning, can generate thousands to tens of thousands of quantitative phenotypes from a single cohort. In such settings, running genome-wide association…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-24 Xingzhong Zhao , Ziqian Xie , Islam , Sheikh Muhammad Saiful , Tian Xia , Chen , Cheng , Degui Zhi

This paper provides a global picture about the deployment of networked processing services for genomic data sets. Many current research make an extensive use genomic data, which are massive and rapidly increasing over time. They are…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-09-27 Gianluca Reali , Mauro Femminella , Emilia Nunzi , Dario Valocchi

The evolution of the Internet and computer applications have generated colossal amount of data. They are referred to as Big Data and they consist of huge volume, high velocity, and variable datasets that need to be managed at the right…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-08-13 Youssef Bassil

Federated learning (FL) is emerging as a privacy-aware alternative to classical cloud-based machine learning. In FL, the sensitive data remains in data silos and only aggregated parameters are exchanged. Hospitals and research institutions…

Machine Learning · Computer Science 2022-05-25 Anne Hartebrodt , Richard Röttger , David B. Blumenthal

Genome Wide Association Studies (GWAS) are used to identify statistically significant genetic variants in case-control studies. GWAS typically use a p-value threshold of 5 x 10-8 to identify highly ranked single nucleotide polymorphisms…

Computational Engineering, Finance, and Science · Computer Science 2018-01-10 Paul Fergus , Casimiro Curbelo Montanez , Basma Abdulaimma , Paulo Lisboa , Carl Chalmers

In this era of large-scale data, distributed systems built on top of clusters of commodity hardware provide cheap and reliable storage and scalable processing of massive data. Here, we review recent work on developing and implementing…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-07-28 Jiyan Yang , Xiangrui Meng , Michael W. Mahoney

This thesis investigates the use of problem-specific knowledge to enhance a genetic algorithm approach to multiple-choice optimisation problems.It shows that such information can significantly enhance performance, but that the choice of…

Neural and Evolutionary Computing · Computer Science 2010-07-05 Uwe Aickelin

In recent years, ideas from statistics and scientific computing have begun to interact in increasingly sophisticated and fruitful ways with ideas from computer science and the theory of algorithms to aid in the development of improved…

Data Structures and Algorithms · Computer Science 2010-10-11 Michael W. Mahoney