相关论文: On a chain of fragmentation equations for duplicat…
An explicit solution for a general two-type birth-death branching process with one way mutation is presented. This continuous time process mimics the evolution of resistance to treatment, or the onset of an extra driver mutation during…
We study a stochastic model based on a modified fragmentation of a finite interval. The mechanism consists in cutting the interval at a random location and substituting a unique fragment on the right of the cut to regenerate and preserve…
Since the sequencing of large genomes, many statistical features of their sequences have been found. One intriguing feature is that certain subsequences are much more abundant than others. In fact, abundances of subsequences of a given…
Several processes in the cell, such as gene regulation, start when key proteins recognise and bind to short DNA sequences. However, as these sequences can be hundreds of million times shorter than the genome, they are hard to find by simple…
The analysis of correlations of amino acid occurrences in globular proteins has led to the development of statistical tools that can identify native contacts -- portions of the chains that come to close distance in folded structural…
We propose a stochastic model of a fragmentation process, developed by taking into account fragment lifetime as a function of their size based on the Gibrat process. If lifetime is determined by a power function of fragment size, numerical…
Sequencing by synthesis is used in many next-generation DNA sequencing technologies. Some of the technologies, especially those exploring the principle of single-molecule sequencing, allow incomplete nucleotide incorporation in each cycle.…
Some natural proteins display recurrent structural patterns. Despite being highly similar at the tertiary structure level, repetitions within a single repeat protein can be extremely variable at the sequence level. We propose a mathematical…
In segmentation problems, inference on change-point position and model selection are two difficult issues due to the discrete nature of change-points. In a Bayesian context, we derive exact, non-asymptotic, explicit and tractable formulae…
The variation in DNA copy number carries information on the modalities of genome evolution and misregulation of DNA replication in cancer cells; its study can be helpful to localize tumor suppressor genes, distinguish different populations…
Based on a model first studied in [Physical Review A, vol. 43, 5240-5260,1991] properties of correlation function for expansion-modification systems are developed. The existence of several characteristic exponents is proved. The…
It is a well-known fact that genetic sequences may contain sections with repeated units, called repeats, that differ in length over a population, with a length distribution of geometric type. A simple class of recombination models with…
The evolution of the full repertoire of proteins encoded in a given genome is mostly driven by gene duplications, deletions, and sequence modifications of existing proteins. Indirect information about relative rates and other intrinsic…
With the number of sequenced genomes now over one hundred, and the availability of rough functional annotations for a substantial proportion of their genes, it has become possible to study the statistics of gene content across genomes. Here…
Power law distributions have been repeatedly observed in a wide variety of socioeconomic, biological and technological areas. In many of the observations, e.g., city populations and sizes of living organisms, the objects of interest evolve…
SAXS studies of four 60 base-pair DNA duplexes with sequences closely related to part of the GAGE6 (G-antigen 6) promoter have been performed to study the role of DNA conformations in solution and their potential relationship to DNA-protein…
We discuss a model of protein conformations where the conformations are combinations of short fragments from some small set. For these fragments we consider a distribution of frequencies of occurrence of pairs (sequence of amino acids,…
We build networks of genetic similarity in which the nodes are organisms sampled from biological populations. The procedure is illustrated by constructing networks from genetic data of a marine clonal plant. An important feature in the…
Although cumulative family name distributions in many countries exhibit power-law forms, there also exist counterexamples. The origin of different family name distributions across countries is discussed analytically in the framework of a…
Helical molecules change their twist number under the effect of a mechanical load. We study the twist-stretch relation for a set of short DNA molecules modeled by a mesoscopic Hamiltonian. Finite temperature path integral techniques are…