相关论文: On the Hypercube Structure of the Genetic Code
In this study, the distributions of protein structure classes (or folding types) of experimentally determined structures from a legacy dataset and a comprehensive database (SCOP) are modeled precisely with geometric constructs such as…
Natural protein sequences that self-assemble to form globular structures are compact with high packing densities in the folded states. It is known that proteins unfold upon addition of denaturants, adopting random coil structures. The…
In the framework of the crystal basis model of the genetic code, where each codon is assigned to an irreducible representation of $U_{q \to 0}(sl(2) \oplus sl(2))$, single base mutation matrices are introduced. The strength of the mutation…
Each human genome is a 3 billion base pair set of encoding instructions. Decoding the genome using deep learning fundamentally differs from most tasks, as we do not know the full structure of the data and therefore cannot design…
We revisit the notion of gene regulatory code in embryonic development in the light of recent findings about genome spatial organisation. By analogy with the genetic code, we posit that the concept of code can only be used if the…
This letter reports complete sets of two-fold symmetries between partitions of the universal genetic code. By substituting bases at each position of the codons according to a fixed rule, it happens that properties of the degeneracy pattern…
Coding information is the main source of heterogeneity (non-randomness) in the sequences of bacterial genomes. This information can be naturally modeled by analysing cluster structures in the "in-phase" triplet distributions of relatively…
Protein design is a technique to engineer proteins by modifying their sequence to obtain novel functionalities. In this method, amino acids in the sequence are permutated to find the low energy states satisfying the configuration. However,…
An In Silico model to relate the properties of proteins to the structure, sequence, function and evolutionary history of proteins is shown. The derived ideal sequences for amino acid residues in proteins can then be considered as attractors…
DNA storage has emerged as an important area of research. The reliability of DNA storage system depends on designing the DNA strings (called DNA codes) that are sufficiently dissimilar. In this work, we introduce DNA codes that satisfy a…
The sequence of a protein is not only constrained by its physical and biochemical properties under current selection, but also by features of its past evolutionary history. Understanding the extent and the form that these evolutionary…
The genetic code underlying protein synthesis is a canonical example of a degenerate biological system. Degeneracies in physical and biological systems can be lifted by external perturbations thus allowing degenerate systems to exhibit a…
The problem of the directionality of genome evolution is studied from the information-theoretic view. We propose that the function-coding information quantity of a genome always grows in the course of evolution through sequence duplication,…
Biologists have long sought a way to explain how statistical properties of genetic sequences emerged and are maintained through evolution. On the one hand, non-random structures at different scales indicate a complex genome organisation. On…
Proteins are large biomolecules that regulate all living organisms and consist of one or several chains. The primary structure of a protein chain is a sequence of amino acid residues whose three main atoms (alpha-carbon, nitrogen, and…
Because of the double-helical structure of DNA, in which two strands of complementary nucleotides intertwine around each other, a covalently closed DNA molecule with no interruptions in either strand can be viewed as two interlocked…
How DNA is mapped to functional proteins is a basic question of living matter. We introduce and study a physical model of protein evolution which suggests a mechanical basis for this map. Many proteins rely on large-scale motion to…
Two new constructions are presented for coils and snakes in the hypercube. Improvements are made on the best known results for snake-in-the-box coils of dimensions 9, 10 and 11, and for some other circuit codes of dimensions between 8 and…
A genetic algorithm is suitable for exploring large search spaces as it finds an approximate solution. Because of this advantage, genetic algorithm is effective in exploring vast and unknown space such as molecular search space. Though the…
During the course of evolution, an organism's genome can undergo changes that affect the large-scale structure of the genome. These changes include gene gain, loss, duplication, chromosome fusion, fission, and rearrangement. When gene gain…