Genomics
Modern high throughput sequencing technologies like metagenomic sequencing generate millions of sequences which have to be classified based on their taxonomic rank. Modern approaches either apply local alignment and comparison to existing…
Endophytes are tiny organisms present in living tissues of distinct plants and have been extensively studied for their endophytic microbial complement. Roots of Rosmarinus officinalis were subjected to the isolation of endophytic fungi and…
Summary: Mapping ancient DNA to a reference genome is challenging as it involves numerous steps, is time-consuming and has to be repeated within a study to assess the quality of extracts and libraries; as a result, the mapping needs to be…
Autism spectrum disorder (ASD) is a heterogeneous neurodevelopmental disorder (NDD) that is caused by genetic, epigenetic, and environmental factors. Recent advances in genomic analysis have uncovered numerous candidate genes with common…
Spatially annotated single-cell datasets provide unprecedented opportunities to dissect cell-cell communication in development and disease. Heterotypic signaling includes interactions between different cell types and is well established in…
We present several algorithms designed to learn a pattern of correspondence between two data sets in situations where it is desirable to match elements that exhibit a relationship belonging to a known parametric model. In the motivating…
Motivation: 3D structures of proteins provide rich information for understanding their biochemical roles. Identifying the representative protein structures for protein sequences is essential for analysis of proteins at proteome scale.…
Escherichia coli is one of many bacterial inhabitants found in human intestines and any adaptation as a result of mutations may affect its host. A commonly used technique employed to study these mutations is Restriction Fragment Length…
An expression quantitative trait locus (eQTL) is a chromosomal region where genetic variants are associated with the expression levels of certain genes that can be both nearby or distant. The identifications of eQTLs for different tissues,…
The study of common heritability, or co-heritability, among multiple traits has been widely established in quantitative and molecular genetics. However, in bacteria, genome-based estimation of heritability has only been considered very…
Latent variable models such as the Variational Auto-Encoder (VAE) have become a go-to tool for analyzing biological data, especially in the field of single-cell genomics. One remaining challenge is the interpretability of latent variables…
The problem of estimating the structure of a graph from observed data is of growing interest in the context of high-throughput genomic data, and single-cell RNA sequencing in particular. These, however, are challenging applications, since…
Various machine-learning models, including deep neural network models, have already been developed to predict deleteriousness of missense (non-synonymous) mutations. Potential improvements to the current state of the art, however, may still…
Leukemia, the cancer of blood cells, originates in the blood-forming cells of the bone marrow. In Chronic Myeloid Leukemia (CML) conditions, the cells partially become mature that look like normal white blood cells but do not resist…
As the availability of omics data has increased in the last few years, more multi-omics data have been generated, that is, high-dimensional molecular data consisting of several types such as genomic, transcriptomic, or proteomic data, all…
Biological systems can be studied at multiple levels of information, including gene, protein, RNA and different interaction networks levels. The goal of this work is to explore how the fusion of systems' level information with temporal…
A core task in precision agriculture is the identification of climatic and ecological conditions that are advantageous for a given crop. The most succinct approach is geolocation, which is concerned with locating the native region of a…
Motivation: As viruses that mainly infect bacteria, phages are key players across a wide range of ecosystems. Analyzing phage proteins is indispensable for understanding phages' functions and roles in microbiomes. High-throughput sequencing…
Motivation: Detection of structural variants (SV) from the alignment of sample DNA reads to the reference genome is an important problem in understanding human diseases. Long reads that can span repeat regions, along with an accurate…
Classification algorithms using RNA-Sequencing (RNA-Seq) data as input are used in a variety of biological applications. By nature, RNA-Seq data is subject to uncontrolled fluctuations both within and especially across datasets, which…