Related papers: Multifractal characterisation of complete genomes
This paper introduces the notion of measure representation of DNA sequences. Spectral analysis and multifractal analysis are then performed on the measure representations of a large number of complete genomes. The main aim of this paper is…
The coding and noncoding length sequences constructed from a complete genome are characterised by multifractal analysis. The dimension spectrum $D_{q}$ and its derivative, the 'analogous' specific heat $C_{q}$, are calculated for the coding…
This paper considers the problem of matching fragment to organism using its complete genome. Our method is based on the probability measure representation of a genome. We first demonstrate that these probability measures can be modelled as…
Under the formalism of annealed averaging of the partition function, a type of random multifractal measures with their multipliers satisfying exponentially distributed is investigated in detail. Branching emerges in the curve of generalized…
The characterization of the long-range order and fractal properties of DNA sequences has proved a difficult though highly rewarding task due mainly to the mosaic character of DNA consisting of many interwoven patches of various lengths with…
Herein it is shown that in order to study the statistical properties of DNA sequences in bacterial chromosomes it suffices to consider only one half of the chromosome because they are similar to its corresponding complementary sequence in…
Under the formalism of annealed averaging of the partition function, two types of random multifractal measures with their probability of multipliers satisfying power distribution and triangular distribution are investigated mathematically.…
This is a review of a set of recent papers with some new data added. After a brief biological introduction a visualization scheme of the string composition of long DNA sequences, in particular, of bacterial complete genomes, will be…
This paper presents a probabilistic approach for DNA sequence analysis. A DNA sequence consists of an arrangement of the four nucleotides A, C, T and G and different representation schemes are presented according to a probability measure…
We present a novel method for determining multi-fractal properties from experimental data. It is based on maximising the likelihood that the given finite data set comes from a particular set of parameters in a multi-parameter family of well…
We show that textual analysis of microbial genomes reveal telling footprints of the early evolution of the genomes. The frequencies of word occurrence of random DNA sequences considered as texts in their four nucleotides are expected to…
We introduce a new approach to constructing networks with realistic features. Our method, in spite of its conceptual simplicity (it has only two parameters) is capable of generating a wide variety of network types with prescribed…
Large-scale dynamical properties of complete chromosome DNA sequences of eukaryotes are considered. By the proposed deterministic models with intermittency and symbolic dynamics we describe a wide spectrum of large-scale patterns inherent…
The so called long range correlation properties of DNA sequences are studied using the variance analyses of the density distribution of a single or a group of nucleotides in a model independent way. This new method which was suggested…
We characterise probability distributions via a martingale property associated with a natural generalisation of record values, known as $\delta$-records. For an independent and identically distributed sequence $(X_n)$ with running maximum…
A gene expression compendium is a heterogeneous collection of gene expression experiments assembled from data collected for diverse purposes. The widely varied experimental conditions and genetic backgrounds across samples creates a…
Extended geometric distribution is defined and its mixture is characterized by the property of having completely monotone probability sequence. Also, convolution equations and probability generating functions are used to characterize…
Sequencing by synthesis is used in many next-generation DNA sequencing technologies. Some of the technologies, especially those exploring the principle of single-molecule sequencing, allow incomplete nucleotide incorporation in each cycle.…
Application of the exact statistical inference frequently leads to a non-standard probability distributions of the considered estimators or test statistics. The exact distributions of many estimators and test statistics can be specified by…
The main statistical distributions applicable to the analysis of genome architecture and genome tracks are briefly discussed and critically assessed. Although the observed features in distributions of element lengths can be equally well…