English
Related papers

Related papers: Fast computation of kernel statistics using genoty…

200 papers

The integration of knowledge graphs and graph machine learning (GML) in genomic data analysis offers several opportunities for understanding complex genetic relationships, especially at the RNA level. We present a comprehensive approach for…

Artificial Intelligence · Computer Science 2024-08-06 Shivika Prasanna , Ajay Kumar , Deepthi Rao , Eduardo Simoes , Praveen Rao

Since most analysis software for genome-wide association studies (GWAS) currently exploit only unrelated individuals, there is a need for efficient applications that can handle general pedigree data or mixtures of both population and…

Applications · Statistics 2014-12-23 Hua Zhou , John Blangero , Thomas D. Dyer , Kei-hang K. Chan , Kenneth Lange , Eric M. Sobel

Substantial progress has been made in identifying single genetic variants predisposing to common complex diseases. Nonetheless, the genetic etiology of human diseases remains largely unknown. Human complex diseases are likely influenced by…

Methodology · Statistics 2014-05-27 Zihuai He , Min Zhang , Xiaowei Zhan , Qing Lu

Sequencing-based studies are emerging as a major tool for genetic association studies of complex diseases. These studies pose great challenges to the traditional statistical methods (e.g., single-locus analyses based on regression methods)…

Methodology · Statistics 2025-08-18 Changshuai Wei , Qing Lu

A key challenge in genomics is to identify genetic variants that distinguish patients with different survival time following diagnosis or treatment. While the log-rank test is widely used for this purpose, nearly all implementations of the…

Quantitative Methods · Quantitative Biology 2013-09-18 Fabio Vandin , Alexandra Papoutsaki , Benjamin J. Raphael , Eli Upfal

Genome wide association studies directly assay 10^6 single nucleotide polymorphisms (SNPs) across a study cohort. Probabilistic estimation of additional sites by genotype imputation can increase this set of variants by 10- to 40-fold. Even…

Quantitative Methods · Quantitative Biology 2013-11-19 Cameron Palmer , Itsik Pe'er

After the completion of human genome sequence was anounced, it is evident that interpretation of DNA sequences is an immediate task to work on. For understanding their signals, improvement of present sequence analysis tools and developing…

Computational Complexity · Computer Science 2007-05-23 Gene Kim , MyungHo Kim

Dimensionality reduction techniques are essential for visualizing and analyzing high-dimensional biological sequencing data. t-distributed Stochastic Neighbor Embedding (t-SNE) is widely used for this purpose, traditionally employing the…

Machine Learning · Computer Science 2025-12-19 Avais Jan , Prakash Chourasia , Sarwan Ali , Murray Patterson

Genomic regions (or loci) displaying outstanding correlation with some environmental variables are likely to be under selection and this is the rationale of recent methods of identifying selected loci and retrieving functional information…

Populations and Evolution · Quantitative Biology 2013-08-13 Gilles Guillot

Meta-analysis of multiple genome-wide association studies (GWAS) is effective for detecting single or multi marker associations with complex traits. We develop a flexible procedure ("STAMP") based on mixture models to perform region based…

Methodology · Statistics 2018-01-01 Andriy Derkach , Ruth M. Pfeiffer

Summary: BGT is a compact format, a fast command line tool and a simple web application for efficient and convenient query of whole-genome genotypes and frequencies across tens to hundreds of thousands of samples. On real data, it encodes…

Genomics · Quantitative Biology 2017-08-07 Heng Li

We present an alternative method for genome-wide association studies (GWAS) that is more powerful than the regular GWAS method for locus detection. The regular GWAS method suffers from a substantial multiple-testing burden because of the…

Applications · Statistics 2018-12-19 William Denault , Håkon K. Gjessing , Julius Juodakis , Bo Jacobsson , Astanand Jugessur

The computational problem of inferring the full haplotype of a cell starting from read sequencing data is known as haplotype assembly, and consists in assigning all heterozygous Single Nucleotide Polymorphisms (SNPs) to exactly one of the…

Representation learning is an important step in the machine learning pipeline. Given the current biological sequencing data volume, learning an explicit representation is prohibitive due to the dimensionality of the resulting feature…

Machine Learning · Computer Science 2023-04-04 Sarwan Ali , Usama Sardar , Murray Patterson , Imdad Ullah Khan

Motivation: Quality control of genomic data is an essential but complicated multi-step procedure, often requiring separate installation and expert familiarity with a combination of disparate bioinformatics tools. Results: To provide an…

Genomics · Quantitative Biology 2021-05-06 Christina Vasilopoulou , Benjamin Wingfield , Andrew P. Morris , William Duddy

Studying the effects of groups of Single Nucleotide Polymorphisms (SNPs), as in a gene, genetic pathway, or network, can provide novel insight into complex diseases, above that which can be gleaned from studying SNPs individually. Common…

Applications · Statistics 2017-10-12 Ryan Sun , Xihong Lin

Efficiently computing a subset of a correlation matrix consisting of values above a specified threshold is important to many practical applications. Real-world problems in genomics, machine learning, finance other applications can produce…

Computation · Statistics 2016-03-15 James Baglama , Michael Kane , Bryan Lewis , Alex Poliakov

The graphlet kernel is a classical method in graph classification. It however suffers from a high computation cost due to the isomorphism test it includes. As a generic proxy, and in general at the cost of losing some information, this test…

Machine Learning · Computer Science 2020-10-19 Hashem Ghanem , Nicolas Keriven , Nicolas Tremblay

Fast and cheaper next generation sequencing technologies will generate unprecedentedly massive and highly-dimensional genomic and epigenomic variation data. In the near future, a routine part of medical record will include the sequenced…

Genomics · Quantitative Biology 2013-01-17 Momiao Xiong , Long Ma

Kernel methods are a highly effective and widely used collection of modern machine learning algorithms. A fundamental limitation of virtually all such methods are computations involving the kernel matrix that naively scale quadratically…

Machine Learning · Computer Science 2021-06-09 John Paul Ryan , Sebastian Ament , Carla P. Gomes , Anil Damle