Related papers: The Triplet Genetic Code had a Doublet Predecessor
The primitive data for deducing the Miyazawa-Jernigan contact energy or BLOSUM score matrix consists of pair frequency counts. Each amino acid corresponds to a conditional probability distribution. Based on the deviation of such conditional…
All known terrestrial proteins are coded as continuous strings of ~20 amino acids. The patterns formed by the repetitions of elements in groups of finite sequences describes the natural architectures of protein families. We present a method…
The set of known dialects of the genetic code (GC) is analyzed from the viewpoint of the genetic octave Yin-Yang-algebra. This algebra was described in the previous author's publications. The algebra was discovered on the basis of…
We present a geometrical analysis of the protrusion statistics of side chains in more than 4,000 high-resolution protein structures. We employ a coarse-grained representation of the protein backbone viewed as a linear chain of C{\alpha}…
The coding space of protein sequences is shaped by evolutionary constraints set by requirements of function and stability. We show that the coding space of a given protein family--the total number of sequences in that family--can be…
We showed in this paper that similarity network can be used as an powerful tools to study the relationship of tRNA genes. We constructed a network of 3719 tRNA gene sequences using simplest alignment and studied its topology, degree…
We present a completely new version of our arithmetic model of the standard genetic code and compute in a straightforward manner the exact numeric degeneracies of the five multiplets without any trick for the doublets and the sextets, as we…
The primary structure of a ribonucleic acid (RNA) molecule can be represented as a sequence of nucleotides (bases) over the alphabet {A, C, G, U}. The secondary or tertiary structure of an RNA is a set of base pairs which form bonds between…
Animal mitochondrial genomes usually have two transfer RNAs for Leucine: one, with anticodon UAG, translates the four-codon family CUN, whilst the other, with anticodon UAA, translates the two-codon family UUR. These two genes must differ…
In several recent papers new gene-detection algorithms were proposed for detecting protein-coding regions without requiring learning dataset of already known genes. The fact that unsupervised gene-detection is possible closely connected to…
A new version of DNA walks, where nucleotides are regarded unequal in their contribution to a walk is introduced, which allows us to study thoroughly the "fine structure" of nucleotide sequences. The approach is based on the assumption that…
We propose a partitioning of the set of unlabelled, connected cubic graphs into two disjoint subsets named genes and descendants, where the cardinality of the descendants is much larger than that of the genes. The key distinction between…
Coding information is the main source of heterogeneity (non-randomness) in the sequences of bacterial genomes. This information can be naturally modeled by analysing cluster structures in the "in-phase" triplet distributions of relatively…
Eukaryote genomes contain excessively introns, inter-genic and other non-genic sequences that appear to have no vital functional role or phenotype manifestation. Their existence, a long-standing puzzle, is viewed from the principle of…
Statistical analysis of distributions of occurrence frequencies of short words in 108 microbial complete genomes reveals the existence of a set of universal "root-sequence lengths" shared by all microbial genomes. These lengths and their…
The GC-content is very variable in different genome regions and species but although many hypothesis we still do not know the reason why. Here we show that a relationship exists with the mutation rate, in particular we noticed a new…
The base sequences of DNA contain the genetic code and to decode it a double helical DNA has to open its base pairs. Recent studies have shown that one can use a third strand to identify the base sequences without opening the double helix…
Complex systems with tightly coadapted parts frequently appear in living systems and are difficult to account for through Darwinian evolution, that is random variation and natural selection, if the constituent parts are independently coded…
In sexual population, recombination reshuffles genetic variation and produces novel combinations of existing alleles, while selection amplifies the fittest genotypes in the population. If recombination is more rapid than selection,…
A heuristic diagram of the evolution of the standard genetic code is presented. It incorporates, in a way that resembles the energy levels of an atom, the physical notion of broken symmetry and it is consistent with original ideas by Crick…