Related papers: The combinatorics of overlapping genes
The simplest possible informational heteropolymer requires only a two-letter alphabet to be able to store information. The evolutionary choice of four monomers in the informational biomolecules RNA/DNA or their progenitors is intriguing,…
We have presented the basic knowledge on the structure of molecules coding the genetic information, mechanisms of transfer of this information from DNA to proteins and phenomena connected with replication of DNA. In particular, we have…
The idea of the evolution of the genetic code from the CG to the CGUA alphabet has been developed further. The assumption of the originally triplet structure of the genetic code has been substantiated. The hypothesis of the emergence of…
Inferring the structural properties of a protein from its amino acid sequence is a challenging yet important problem in biology. Structures are not known for the vast majority of protein sequences, but structure is critical for…
We propose a stochastic model for gene transcription coupled to DNA supercoiling, where we incorporate the experimental observation that polymerases create supercoiling as they unwind the DNA helix, and that these enzymes bind more…
An RNA sequence is a word over an alphabet on four elements $\{A,C,G,U\}$ called bases. RNA sequences fold into secondary structures where some bases match one another while others remain unpaired. Pseudoknot-free secondary structures can…
Delays in protein synthesis cause a confounding effect when constructing Gene Regulatory Networks (GRNs) from RNA-sequencing time-series data. Accurate GRNs can be very insightful when modelling development, disease pathways, and drug…
Tandem duplication in DNA is the process of inserting a copy of a segment of DNA adjacent to the original position. Motivated by applications that store data in living organisms, Jain {\em et al.} (2016) proposed the study of codes that…
We live in a period where bio-informatics is rapidly expanding, a significant quantity of genomic data has been produced as a result of the advancement of high-throughput genome sequencing technology, raising concerns about the costs…
Functions of chemical composition are complex and discrete in nature making it impossible to optimize them with gradient methods. Genetic algorithms, which do not use derivative information, are used to maximize the thermal conductivity of…
Detecting the interactions of genetic compounds like genes, SNPs, proteins, metabolites, etc. can potentially unravel the mechanisms behind complex traits and common genetic disorders. Several methods have been taken into consideration for…
The majority of the human genome consists of repeated sequences. An important type of repeated sequences common in the human genome are tandem repeats, where identical copies appear next to each other. For example, in the sequence…
Higher-dimensional rewriting is founded on a duality of rewrite systems and cell complexes, connecting computational mathematics to higher categories and homotopy theory: the two sides of a rewrite rule are two halves of the boundary of an…
Protein sequences serve as a natural record of the evolutionary constraints that shape their functional structures. We show that it is possible to use only sequence information to go beyond predicting native structures and global stability…
We consider a model of two (fully) compact polymer chains, coupled through an attractive interaction. These compact chains are represented by Hamiltonian paths (HP), and the coupling favors the existence of common bonds between the chains.…
For much of biology, the manner in which genotype maps to phenotype remains a fundamental mystery. The few maps that are known tend to show modular pleiotropy: sets of phenotypes are determined by distinct sets of genes. One key map that…
In a certain way, this paper presents the continuation of the previous one which discussed the harmonic structure of the genetic code (Rakocevic, 2004). Several new harmonic structures presented in this paper, through specific unity and…
This article presents a theoretical investigation of generalized encoded forms of networks in a uniform multidimensional space. First, we study encoded networks with (finite) arbitrary node dimensions (or aspects), such as time instants or…
It has been proposed that the degeneracy of the genetic code,i.e., the phenomenon that different codons (base triplets) of DNA are transcribed into the same amino acid, may be interpreted as the result of a symmetry breaking process. In the…
The repeat content and heterozygosity rate of a target genome are important factors in determining the feasibility of achieving a complete telomere-to-telomere assembly. The mathematical relationship between the required coverage and read…