Genomics
Functional or non-coding RNAs are attracting more attention as they are now potentially considered valuable resources in the development of new drugs intended to cure several human diseases. The identification of drugs targeting the…
To date, twelve complete genomes representing eleven species belonging to six genera have been sequenced in salmonids. For the genus Salvelinus, it was supposed to sequence the genome of Arctic char, one of the most variable species of…
Nanopore sequencing technology has the potential to render other sequencing technologies obsolete with its ability to generate long reads and provide portability. However, high error rates of the technology pose a challenge while generating…
The human T cell repertoire is generated by the rearrangement of variable (V), diversity (D) and joining (J) segments on the T cell receptor (TCR) loci. To determine whether the structural ordering of these gene segments on the TCR loci…
A blood cell lineage consists of several consecutive developmental stages from the pluripotent or multipotent stem cell to a particular stage of terminally differentiated cells. There is considerable interest in identifying the key…
Local ancestry inference (LAI) allows identification of the ancestry of all chromosomal segments in admixed individuals, and it is a critical step in the analysis of human genomes with applications from pharmacogenomics and precision…
The majority of cancer treatments end in failure due to Intra-Tumor Heterogeneity (ITH). ITH in cancer is represented by clonal evolution where different sub-clones compete with each other for resources under conditions of Darwinian natural…
The aim of this work is to provide a rigorous mathematical analysis of a stochastic concatenation model presented by Sobottka and Hart (2011) which allows approximation of the first-order stochastic structure in bacterial DNA by means of a…
Finding tumour genetic markers is essential to biomedicine due to their relevance for cancer detection and therapy development. In this paper, we explore a recently released dataset of chromosome rearrangements in 2,586 cancer patients,…
Explicit accounting for copy number alterations can dramatically improve mutation frequency estimates, leading to more accurate phylogeny reconstructions and subclone characterizations.
Genome-wide association studies (GWAS) provide a means of examining the common genetic variation underlying a range of traits and disorders. In addition, it is hoped that GWAS may provide a means of differentiating affected from unaffected…
Protein domains are highly conserved functional units of proteins. Because they carry functionally significant information, the majority of the coding disease variants are located on domains. Additionally, domains are specific units of the…
Determining the primary site of origin for metastatic tumors is one of the open problems in cancer care because the efficacy of treatment often depends on the cancer tissue of origin. Classification methods that can leverage tumor genomic…
Next Generation Sequencing can sample the whole genome (WGS) or the 1-2% of the genome that codes for proteins called the whole exome (WES). Machine learning approaches to variant calling achieve high accuracy in WGS data, but the reduced…
Objectives Lung squamous cell carcinoma (LUSC) often diagnosed as advanced with poor prognosis. The mechanisms of its pathogenesis and prognosis require urgent elucidation. This study was performed to screen potential biomarkers related to…
Drug development is a very costly and lengthy process, while repositioned or repurposed drugs could be brought into clinical practice within a shorter time-frame and at a much reduced cost. The past decade has observed a massive growth in…
Reconstructing components of a genomic mixture from data obtained by means of DNA sequencing is a challenging problem encountered in a variety of applications including single individual haplotyping and studies of viral communities.…
We revisit the notion of gene regulatory code in embryonic development in the light of recent findings about genome spatial organisation. By analogy with the genetic code, we posit that the concept of code can only be used if the…
Biomedical data, particularly in the field of genomics, has characteristics which make it challenging for machine learning applications - it can be sparse, high dimensional and noisy. Biomedical applications also present challenges to model…
With recent advances in sequencing technologies, large amounts of epigenomic data have become available and computational methods are contributing significantly to the progress of epigenetic research. As an orthogonal approach to methods…