Related papers: Genetic Code: Four Diversity Types of Protein Amin…
What are proteins made from, as the working parts of the living cells protein machines? To answer this question, we need a technology to disassemble proteins onto elementary func-tional details and to prepare lumped description of such…
The genesis of the stand genetic code is considered as a result of a fusion of two AU- and GC-codes distributed in two dominant and two recessive domains. The fusion of these codes is described with simple empirical rules. This formal…
This paper reports about an approach to the classification of proteins' primary structures taking advantage of the Self Organizing Maps algorithm and of a numerical coding of the aminoacids based upon their physico-chemical properties.…
Protein structure prediction is a challenging and unsolved problem in computer science. Proteins are the sequence of amino acids connected together by single peptide bond. The combinations of the twenty primary amino acids are the…
New analyses of the organization of the genetic code system together with their relation to the two classes of aminoacyl-tRNA synthetases are reported in this work. A closer inspection revealed how the enzymes and the 20 amino acids of the…
We present a structural data set of the 20 proteinogenic amino acids and their amino-methylated and acetylated (capped) dipeptides. Different protonation states of the backbone (uncharged and zwitterionic) were considered for the amino…
Atomic packing is an important metric for characterizing protein structures, as it significantly influences various features including the stability, the rate of evolution and the functional roles of proteins. Packing in protein structures…
In this Communication we present statistical analysis of conservation profiles in families of homologous sequences for nine proteins whose folding nucleus was determined by protein engineering methods. We show that in all but one protein…
Evolution consists of distinct stages: cosmological, biological, linguistic. Since biology verges on natural sciences and linguistics, we expect that it shares structures and features from both forms of knowledge. Indeed, in DNA we…
Recently, the bond lengths of the molecular components of nucleic acids and of caffeine and related molecules were shown to be sums of the appropriate covalent radii of the adjacent atoms. Thus, each atom was shown to have its specific…
Proteins have evolved through mutations, amino acid substitutions, since life appeared on Earth, some 109 years ago. The study of these phenomena has been of particular significance because of their impact on protein stability, function,…
In this short paper, it is shown that the multiplet structure of the standard genetic code is derivable from the total number of nucleotides contained in 64 codons, 192, a small number. The degeneracy class-number is derived as the number…
We perform an exhaustive analysis of genome statistics for organisms, particularly extremophiles, growing in a wide range of physicochemical conditions. Specifically, we demonstrate how the correlation between the frequency of amino acids…
The biological distinction between the base positions in the codon, the chemical types of bases (purine and pyrimidine) and their hydrogen bond number have been the most relevant codon properties used in the genetic code analysis. Now,…
Background: There is a 3-fold redundancy in the Genetic Code; most amino acids are encoded by more than one codon. These synonymous codons are not used equally; there is a Codon Usage Bias (CUB). This article will provide novel information…
In protein secondary structure prediction, each amino acid in sequence is typically treated as a distinct category and represented by a one-hot vector. In this study, we developed two novel chemical representations for amino acids utilizing…
Most amino acids are encoded by multiple synonymous codons. For an amino acid, some of its synonymous codons are used much more rarely than others. Analyses of positions of such rare codons in protein sequences revealed that rare codons can…
Determining the full complement of protein-coding genes is a key goal of genome annotation. The most powerful approach for confirming protein coding potential is the detection of cellular protein expression through peptide mass spectrometry…
Methods for alignment of protein sequences typically measure similarity by using substitution matrix with scores for all possible exchanges of one amino acid with another. Although widely used, the matrices derived from homologous sequence…
A simple lattice model for proteins that allows for distinct sizes of the amino acids is presented. The model is found to lead to a significant number of conformations that are the unique ground state of one or more sequences or encodable.…