English
Related papers

Related papers: Fast computation of kernel statistics using genoty…

200 papers

Clustering genotypes based upon their phenotypic characteristics is used to obtain diverse sets of parents that are useful in their breeding programs. The Hierarchical Clustering (HC) algorithm is the current standard in clustering of…

Machine Learning · Computer Science 2020-09-22 Aditya A. Shastri , Kapil Ahuja , Milind B. Ratnaparkhe , Yann Busnel

In genetic association studies, detecting phenotype-genotype association is a primary goal. We assume that the relationship between the data -phenotype, genetic markers and environmental covariates - can be modelled by a generalized linear…

Methodology · Statistics 2020-04-13 K. K. Halle , Ø. Bakke , S. Djurovic , A. Bye , E. Ryeng , U. Wisløff , O. A. Andreassen , M. Langaas

There is a growing interest in cell-type-specific analysis from bulk samples with a mixture of different cell types. A critical first step in such analyses is the accurate estimation of cell-type proportions in a bulk sample. Although many…

Methodology · Statistics 2022-09-12 Biao Cai , Jingfei Zhang , Hongyu Li , Chang Su , Hongyu Zhao

Kernel matrices, as well as weighted graphs represented by them, are ubiquitous objects in machine learning, statistics and other related fields. The main drawback of using kernel methods (learning and inference using kernel matrices) is…

Machine Learning · Computer Science 2022-12-02 Ainesh Bakshi , Piotr Indyk , Praneeth Kacham , Sandeep Silwal , Samson Zhou

Clinical adoption of human genome sequencing requires methods with known accuracy of genotype calls at millions or billions of positions across a genome. Previous work showing discordance amongst sequencing methods and algorithms has made…

Genomics · Quantitative Biology 2014-02-18 Justin M. Zook , Brad Chapman , Jason Wang , David Mittelman , Oliver Hofmann , Winston Hide , Marc Salit

This work presents a new approach for classification of genomic sequences from measurements of complex networks and information theory. For this, it is considered the nucleotides, dinucleotides and trinucleotides of a genomic sequence. For…

Computational Engineering, Finance, and Science · Computer Science 2014-12-19 Bruno Mendes Moro Conque , André Yoshiaki Kashiwabara , Fabrício Martins Lopes

The t-Distributed Stochastic Neighbor Embedding (t-SNE) has emerged as a popular dimensionality reduction technique for visualizing high-dimensional data. It computes pairwise similarities between data points by default using an RBF kernel…

Machine Learning · Computer Science 2024-10-22 Sarwan Ali , Prakash Chourasia , Haris Mansoor , Bipin koirala , Murray Patterson

Genome-wide association studies (GWA studies or GWAS) investigate the relationships between genetic variants such as single-nucleotide polymorphisms (SNPs) and individual traits. Recently, incorporating biological priors together with…

Machine Learning · Statistics 2017-09-13 Tao Yang , Paul Thompson , Sihai Zhao , Jieping Ye

Genome-wide Association Studies (GWASs) for complex diseases often collect data on multiple correlated endo-phenotypes. Multivariate analysis of these correlated phenotypes can improve the power to detect genetic variants. Multivariate…

Methodology · Statistics 2015-03-12 Debashree Ray , James S Pankow , Saonli Basu

Approaches for testing sets of variants, such as a set of rare or common variants within a gene or pathway, for association with complex traits are important. In particular, set tests allow for aggregation of weak signal within a set, can…

Genomics · Quantitative Biology 2013-05-28 Jennifer Listgarten , Christoph Lippert , Eun Yong Kang , Jing Xiang , Carl M. Kadie , David Heckerman

Commonly in biomedical research, studies collect data in which an outcome measure contains informative excess zeros; for example when observing the burden of neuritic plaques in brain pathology studies, those who show none contribute to our…

Methodology · Statistics 2018-01-19 Matthew Goodman , Lori Chibnik , Tianxi Cai

One way of investigating how genes affect human traits would be with a genome-wide association study (GWAS). Genetic markers, known as single-nucleotide polymorphism (SNP), are used in GWAS. This raises privacy and security concerns as…

Applications · Statistics 2019-08-02 Jun Jie Sim , Fook Mun Chan , Shibin Chen , Benjamin Hong Meng Tan , Khin Mi Mi Aung

Multi-trait genome-wide association studies (GWAS) use multi-variate statistical methods to identify associations between genetic variants and multiple correlated traits simultaneously, and have higher statistical power than independent…

Genomics · Quantitative Biology 2022-02-10 Muhammad Ammar Malik , Adriaan-Alexander Ludl , Tom Michoel

Quantum Kernel Estimation (QKE) is a technique based on leveraging a quantum computer to estimate a kernel function that is classically difficult to calculate, which is then used by a classical computer for training a Support Vector Machine…

Quantum Physics · Physics 2023-08-01 Marco Russo , Edoardo Giusto , Bartolomeo Montrucchio

Kernel approximation methods create explicit, low-dimensional kernel feature maps to deal with the high computational and memory complexity of standard techniques. This work studies a supervised kernel learning methodology to optimize such…

Machine Learning · Computer Science 2020-02-17 Mert Al , Zejiang Hou , Sun-Yuan Kung

Evaluation and validation of complicated control systems are crucial to guarantee usability and safety. Usually, failure happens in some very rarely encountered situations, but once triggered, the consequence is disastrous. Accelerated…

Machine Learning · Computer Science 2017-10-03 Zhiyuan Huang , Yaohui Guo , Henry Lam , Ding Zhao

Background: Single-cell RNA sequencing (scRNA-seq) yields valuable insights about gene expression and gives critical information about complex tissue cellular composition. In the analysis of single-cell RNA sequencing, the annotations of…

Genomics · Quantitative Biology 2023-03-29 Xiaowen Cao , Li Xing , Elham Majd , Hua He , Junhua Gu , Xuekui Zhang

String Kernel (SK) techniques, especially those using gapped $k$-mers as features (gk), have obtained great success in classifying sequences like DNA, protein, and text. However, the state-of-the-art gk-SK runs extremely slow when we…

Machine Learning · Computer Science 2017-09-19 Ritambhara Singh , Arshdeep Sekhon , Kamran Kowsari , Jack Lanchantin , Beilun Wang , Yanjun Qi

Important objectives in cancer research are the prediction of a patient's risk based on molecular measurements such as gene expression data and the identification of new prognostic biomarkers (e.g. genes). In clinical practice, this is…

Applications · Statistics 2020-04-17 Katrin Madjar , Manuela Zucknick , Katja Ickstadt , Jörg Rahnenführer

Because of the high cost of commercial genotyping chip technologies, many investigations have used a two-stage design for genome-wide association studies, using part of the sample for an initial discovery of ``promising'' SNPs at a less…