Related papers: Statistical analysis of simple repeats in the huma…
Single-nucleotide polymorphisms (SNPs) account for most variations between human genomes. We show how, if the genomes in a database differ only by a reasonable number of SNPs and the substrings between those SNPs are unique, then we can…
The ~4-Mbp basic genome shared by 32 independent isolates of E. coli representing considerable population diversity has been approximated by whole-genome multiple-alignment and computational filtering designed to remove mobile elements and…
We study electronic transport in long DNA chains using the tight-binding approach for a ladder-like model of DNA. We find insulating behavior with localizaton lengths xi ~ 25 in units of average base-pair seperation. Furthermore, we observe…
Intermittent density fluctuations of nucleotide molecules (adenine, guanine, cytosine and thymine) along DNA sequences are studied in the framework of a hierarchical structure (HS) model originally proposed for the study of fully developed…
The macroscopic curvature of double helical DNA induced by regularly repeated adenine tracts is well-known but still puzzling. Its physical origin remains controversial even though it is perhaps the best-documented sequence modulation of…
In this paper, we review the literature on statistical long-range correlation in DNA sequences. We examine the current evidence for these correlations, and conclude that a mixture of many length scales (including some relatively long ones)…
Statistical analysis of distributions of occurrence frequencies of short words in 108 microbial complete genomes reveals the existence of a set of universal "root-sequence lengths" shared by all microbial genomes. These lengths and their…
We consider the task of detecting regulatory elements in the human genome directly from raw DNA. Past work has focused on small snippets of DNA, making it difficult to model long-distance dependencies that arise from DNA's 3-dimensional…
Since the completion of the human genome sequencing project in 2001, significant progress has been made in areas such as gene regulation editing and protein structure prediction. However, given the vast amount of genomic data, the segments…
Nucleic acids are highly deformable helical molecules constantly stretched, twisted and bent in their biological functioning. Single molecule experiments have shown that double stranded (ds)-RNA and standard ds-DNA have opposite…
Splicing sites provide unique statistics in human genome due to their large number and reasonably complete annotation. Analyses of the cumulative SNPs distribution in splicing sites reveal a few interesting observations. While a degree of…
We investigate a densely packed, non-random arrangement of forty-six chromosomes (46,XY) in human nuclei. Here, we model systems-level chromosomal crosstalk by unifying intrinsic parameters (chromosomal length and number of genes) across…
An approximation to the ~4 Mbp basic genome shared by 32 strains of E. coli representing six evolutionary groups has been derived and analyzed computationally. A multiple-alignment of the 32 complete genome sequences was filtered to remove…
The ubiquitous biomacromolecule DNA has an axial rigidity persistence length of ~50 nm, driven by its elegant double helical structure. While double and multiple helix structures appear widely in nature, only rarely are these found in…
Spatial fluctuations of guanine and cytosine base content (GC%) are studied by spectral analysis for the complete set of human genomic DNA sequences. We find that (i) the 1/f^alpha decay is universally observed in the power spectra of all…
Thousands of candidate human-specific regulatory sequences (HSRS) have been identified, supporting the hypothesis that unique to human phenotypes result from human-specific alterations of genomic regulatory networks. Here, conservation…
Motivated by empirical observations of algebraic duplicated sequence length distributions in a broad range of natural genomes, we analytically formulate and solve a class of simple discrete duplication/substitution models that generate…
The evolution of the full repertoire of proteins encoded in a given genome is mostly driven by gene duplications, deletions, and sequence modifications of existing proteins. Indirect information about relative rates and other intrinsic…
A class of nucleosome remodeling motors translocate nucleosomes, to which they are attached, toward the middle of DNA chain in the presence of ATP during in vitro experiments. Such a biological activity is likely based on a physical…
All known terrestrial proteins are coded as continuous strings of ~20 amino acids. The patterns formed by the repetitions of elements in groups of finite sequences describes the natural architectures of protein families. We present a method…