Related papers: A modular Fibonacci sequence in proteins
The systematics of indices of physico-chemical properties of codons and amino acids across the genetic code are examined. Using a simple numerical labelling scheme for nucleic acid bases, data can be fitted as low-order polynomials of the 6…
In sequence-based predictions, conventionally an input sequence is represented by a multiple sequence alignment (MSA) or a representation derived from MSA, such as a position-specific scoring matrix. Recently, inspired by the development in…
The Tribonacci sequence $\mathbb{T}$ is the fixed point of the substitution $\sigma(a,b,c)=(ab,ac,a)$. The main result is twofold: (1) we give the explicit expressions of the numbers of distinct squares and cubes in $\mathbb{T}[1,n]$ (the…
We argue that protein native state structures reside in a novel "phase" of matter which confers on proteins their many amazing characteristics. This phase arises from the common features of all globular proteins and is characterized by a…
This note represents the further progress in understanding the determination of the genetic code by Golden mean (Rakocevic, 1998). Three classes of amino acids that follow from this determination (the 7 "golden" amino acids, 7 of their…
Sequence set is a widely-used type of data source in a large variety of fields. A typical example is protein structure prediction, which takes an multiple sequence alignment (MSA) as input and aims to infer structural information from it.…
The notion of energy landscapes provides conceptual tools for understanding the complexities of protein folding and function. Energy Landscape Theory indicates that it is much easier to find sequences that satisfy the "Principle of Minimal…
For any positive integer $r$, the $r$-Fubini number with parameter $n$, denoted by $F_{n,r}$, is equal to the number of ways that the elements of a set with $n+r$ elements can be weak ordered such that the $r$ least elements are in distinct…
It is well known that every positive integer can be expressed as a sum of nonconsecutive Fibonacci numbers provided the Fibonacci numbers satisfy $F_n =F_{n-1}+F_{n-2}$ for $n\geq 3$, $F_1 =1$ and $F_2 =2$. In this paper, for any…
What are proteins made from, as the working parts of the living cells protein machines? To answer this question, we need a technology to disassemble proteins onto elementary func-tional details and to prepare lumped description of such…
We introduce a new heuristic for the multiple alignment of a set of sequences. The heuristic is based on a set cover of the residue alphabet of the sequences, and also on the determination of a significant set of blocks comprising…
We give necessary and sufficient conditions for a Fibonacci cycle to be residue complete (nondefective). In particular, the Lucas numbers modulo m is residue complete if and only if m = 2,4,6,7,14 or a power of 3.
As the number of solved protein structures increases, the opportunities for meta-analysis of this dataset increase too. Protein structures are known to be formed of domains; structural and functional subunits that are often repeated across…
Motivated by empirical observations of algebraic duplicated sequence length distributions in a broad range of natural genomes, we analytically formulate and solve a class of simple discrete duplication/substitution models that generate…
Proteins are constructed from a limited alphabet of ~20 amino acids, yet the origins and selection of this specific alphabet are unresolved. One largely overlooked aspect is whether elemental composition constrains the range of viable…
Empirical substitution matrices represent the average tendencies of substitutions over various protein families by sacrificing gene-level resolution. We develop a codon-based model, in which mutational tendencies of codon, a genetic code,…
We perform an exhaustive analysis of genome statistics for organisms, particularly extremophiles, growing in a wide range of physicochemical conditions. Specifically, we demonstrate how the correlation between the frequency of amino acids…
Structural classification shows that the number of different protein folds is surprisingly small. It also appears that proteins are built in a modular fashion, from a relatively small number of components. Here we propose to identify the…
A general strategy is described for finding which amino acid sequences have native states in a desired conformation (inverse design). The approach is used to design sequences of 48 hydrophobic and polar aminoacids on three-dimensional…
Residue-residue interactions that fold a protein into a unique three-dimensional structure and make it play a specific function impose structural and functional constraints on each residue site. Selective constraints on residue sites are…