Related papers: An Entropy-Based Technique for Classifying Bacteri…
Background: There is a 3-fold redundancy in the Genetic Code; most amino acids are encoded by more than one codon. These synonymous codons are not used equally; there is a Codon Usage Bias (CUB). This article will provide novel information…
Recent experiments have been able to visualise chromosome organization in fast-growing E.coli cells. However, the mechanism underlying the spatio-temporal organization remains poorly understood. We propose that the DNA adopts a specific…
In the near future, all the human genes will be identified. But understanding the functions coded in the genes is a much harder problem. For example, by using block entropy, one has that the DNA code is closer to a random code then written…
There is a class of entropy-coding methods which do not substitute symbols by code words (such as Huffman coding), but operate on intervals or ranges. This class includes three prominent members: conventional arithmetic coding, range…
Herein it is shown that in order to study the statistical properties of DNA sequences in bacterial chromosomes it suffices to consider only one half of the chromosome because they are similar to its corresponding complementary sequence in…
We propose a method based on finite mixture models for classifying a set of observations into number of different categories. In order to demonstrate the method, we show how the component densities for the mixture model can be derived by…
In special coordinates (codon position--specific nucleotide frequencies) bacterial genomes form two straight lines in 9-dimensional space: one line for eubacterial genomes, another for archaeal genomes. All the 348 distinct bacterial…
Many datasets exhibit a well-defined structure that can be exploited to design faster search tools, but it is not always clear when such acceleration is possible. Here, we introduce a framework for similarity search based on characterizing…
In our paper selected linguistic features of genomes to study the statistics of the gene codes are considered. We present the information theory from which it follows that if the system is described by distributions of hyperbolic type it…
The Chapter starts with introductory information about quantitative linguistics notions, like rank--frequency dependence, Zipf's law, frequency spectra, etc. Similarities in distributions of words in texts with level occupation in quantum…
The entanglement properties of a class of topological stabilizer states, the so called \emph{topological color codes} defined on a two-dimensional lattice or \emph{2-colex}, are calculated. The topological entropy is used to measure the…
Entanglement criteria for general (pure or mixed) states of systems consisting of two identical fermions are introduced. These criteria are based on appropriate inequalities involving the entropy of the global density matrix describing the…
Distribution matching transforms independent and Bernoulli(1/2) distributed input bits into a sequence of output symbols with a desired distribution. Fixed-to-fixed length, invertible, and low complexity encoders and decoders based on…
The principles that govern the organization of genomes, which are needed for a deeper understanding of how chromosomes are packaged and function in eukaryotic cells, could be deciphered if the three-dimensional (3D) structures are known.…
The interplay between bacterial chromosome organization and functions such as transcription and replication can be studied in increasing detail using novel experimental techniques. Interpreting the resulting quantitative data, however, can…
Many proofs in discrete mathematics and theoretical computer science are based on the probabilistic method. To prove the existence of a good object, we pick a random object and show that it is bad with low probability. This method is…
Configurational entropy is an important factor in the free energy change of many macromolecular recognition and binding processes, and has been intensively studied. Despite great progresses that have been made, the global sampling remains…
The presence of clusters of rare codons is known to negatively impact the efficiency and accuracy of protein production. In this paper, we demonstrate a statistical method of identifying such clusters in the coding sequence of a gene. Using…
We present a technique for clustering categorical data by generating many dissimilarity matrices and averaging over them. We begin by demonstrating our technique on low dimensional categorical data and comparing it to several other…
The genetic code underlying protein synthesis is a canonical example of a degenerate biological system. Degeneracies in physical and biological systems can be lifted by external perturbations thus allowing degenerate systems to exhibit a…