Related papers: Determinative degree and nucleotide sequence analy…
The genetic code is the set of rules by which information encoded in genetic material (DNA or RNA sequences) is translated into proteins (amino acid sequences) by living cells. The code defines a mapping between tri-nucleotide sequences,…
A nucleotides sequence is identified, in the two (four) letters alphabet, by the the labels of a vector state of an irreducible representation of U_q(sl(2)) (U_q(sl(2) + sl(2))), in the limit q -> 0. A master equation for the distribution…
Each human genome is a 3 billion base pair set of encoding instructions. Decoding the genome using deep learning fundamentally differs from most tasks, as we do not know the full structure of the data and therefore cannot design…
Inspired by recent successes using single-stranded DNA tiles to produce complex structures, we develop a two-step coarse-graining approach that uses detailed thermodynamic calculations with oxDNA, a nucleotide-based model of DNA, to…
We study the three-dimensional persistent random walk with drift. Then we develop a thermodynamic model that is based on this random walk without assuming the Boltzmann-Gibbs form for the equilibrium distribution. The simplicity of the…
Learning effective protein representations is critical in a variety of tasks in biology such as predicting protein function or structure. Existing approaches usually pretrain protein language models on a large number of unlabeled amino acid…
A representation of the genetic code as a six-dimensional Boolean hypercube is described. This structure is the result of the hierarchical order of the interaction energies of the bases in codon-anticodon recognition. In this paper it is…
The present work is devoted to describe a set of rules explaining the discriminating versus non-discriminating behavior of the di-basic stages and to characterize the role of each base in determining such a behavior. Bases are analyze as…
Understanding the three-dimensional (3D) structure and stability of DNA is fundamental for its biological function and the design of novel drugs. In this study, we introduce an improved coarse-grained (CG) model, incorporating a more…
Selecting appropriate datasets is critical in modern computer vision. However, no general-purpose tools exist to evaluate the extent to which two datasets differ. For this, we propose representing images - and by extension datasets - using…
A new set of DNA base-nucleic acid codes and their hypercomplex number representation have been introduced for taking the probability of each nucleotide into full account. A new scoring system has been proposed to suit the hypercomplex…
Despite the significance of the high flexibility exhibited by short DNAs, there remains an incomplete understanding of their anomalous persistence length. In this study, we propose a novel approach wherein each fundamental characteristic of…
DNA segments and sequences have been studied thoroughly during the past decades. One of the main problems in computational biology is the identification of exon-intron structures inside genes using mathematical techniques. Previous studies…
Understanding the dynamics of genome rearrangements is a major issue of phylogenetics. Phylogenetics is the study of species evolution. A major goal of the field is to establish evolutionary relationships within groups of species, in order…
Problems of search and recognition appear over different scales in biological systems. In this review we focus on the challenges posed by interactions between proteins, in particular transcription factors, and DNA and possible mechanisms…
A self-organizing approach is proposed for gene finding based on the model of codon usage for coding regions and positional preference for noncoding regions. The symmetry between the direct and reverse coding regions is adopted for reducing…
In this letter, we study the structure-transport property relationships of small ligand intercalated DNA molecules using a multiscale modelling approach where extensive ab-initio calculations are performed on numerous MD-simulated…
Genome sequencing technology has improved significantly in few last years and resulted in abundance genetic data. Artificial intelligence has been employed to analyze genetic data in response to its sheer size and variability. Gene…
Phylogenetic networks are a generalization of phylogenetic trees that are used in biology to represent reticulate or non-treelike evolution. Recently, several algorithms have been developed which aim to construct phylogenetic networks from…
While all the information required for the folding of a protein is contained in its amino acid sequence, one has not yet learnt how to extract this information so as to predict the detailed, biological active, three-dimensional structure of…