相关论文: Genetic Code: Four Diversity Types of Protein Amin…
It is widely accepted that (1) the natural or folded state of proteins is a global energy minimum, and (2) in most cases proteins fold to a unique state determined by their amino acid sequence. The H-P (hydrophobic-hydrophilic) model is a…
We examined what determines the designability of 2-letter codes (H and P) lattice proteins from three points of view. First, whether the native structure is searched within all possible structures or within maximally compact structures.…
Despite the variety of protein sizes, shapes, and backbone configurations found in nature, the design of novel protein folds remains an open problem. Within simple lattice models it has been shown that all structures are not equally…
Proteins are constructed from a limited alphabet of ~20 amino acids, yet the origins and selection of this specific alphabet are unresolved. One largely overlooked aspect is whether elemental composition constrains the range of viable…
A two amino acid (hydrophobic and polar) scheme is used to perform the design on target conformations corresponding to the native states of twenty single chain proteins. Strikingly, the percentage of successful identification of the nature…
A new n-dimensional vector space of the DNA sequences on the Galois field of the 64 codons (GF(64)) is proposed. In this vector space gene mutations can be considered linear transformations or translations of the wild type gene. In…
Upon the covalent-bonding hybrid of the nitrogen atoms taken as a measure for the structural regularity in nucleobases, it can be identified that the internal relation within the 20 amino acids follows a cooperative vector-in-space addition…
Evolution of genetic code is studied as the change in the choice of enzymes that are used to synthesize amino acids from the genetic information of nucleic acids. We propose the following theory: the differentiation of physiological states…
Understanding the observed variability in the number of homologs of a gene is a very important, unsolved problem that has broad implications for research into co-evolution of structure and function, gene duplication, pseudogene formation…
In nature the three-dimensional structure of a protein is encoded in the corresponding gene. In this paper we describe a new method for encoding the three-dimensional structure of a protein into a binary sequence. The feature of the method…
This paper presents a probabilistic approach for DNA sequence analysis. A DNA sequence consists of an arrangement of the four nucleotides A, C, T and G and different representation schemes are presented according to a probability measure…
Proteins are the essential drivers of biological processes. At the molecular level, they are chains of amino acids that can be viewed through a linguistic lens where the twenty standard residues serve as an alphabet combining to form a…
Protein structures can be studied as complex networks of interacting amino acids. We study proteins of different structural classes from the network perspective. Our results indicate that proteins, regardless of their structural class, show…
The issues we attempt to tackle here are what the first peptides did look like when they emerged on the primitive earth, and what simple catalytic activities they fulfilled. We conjecture that the early functional peptides were short (3 to…
The primitive data for deducing the Miyazawa-Jernigan contact energy or BLOSUM score matrix consists of pair frequency counts. Each amino acid corresponds to a conditional probability distribution. Based on the deviation of such conditional…
The usage frequencies for codons belonging to quartets are analized, over the whole exonic region, for 92 biological species. Correlation is put into evidence, between the usage frequencies of synonymous codons with third nucleotide A and C…
The intricate three-dimensional geometries of protein tertiary structures underlie protein function and emerge through a folding process from one-dimensional chains of amino acids. The exact spatial sequence and configuration of amino…
We show that a protein can be trained to recognise multiple conformations, analogous to an associative memory, and provide capacity calculations based on energy fluctuations and information theory. Unlike the linear capacity of a Hopfield…
Protein language models learn powerful representations directly from sequences of amino acids. However, they are constrained to generate proteins with only the set of amino acids represented in their vocabulary. In contrast, chemical…
The GC-content is very variable in different genome regions and species but although many hypothesis we still do not know the reason why. Here we show that a relationship exists with the mutation rate, in particular we noticed a new…