Related papers: Genome analyses and modelling the relationship bet…
Probability modelling for DNA sequence evolution is well established and provides a rich framework for understanding genetic variation between samples of individuals from one or more populations. We show that both classical and more recent…
We study the primary DNA structure of four of the most completely sequenced human chromosomes (including chromosome 19 which is the most dense in coding), using Non-extensive Statistics. We show that the exponents governing the decay of the…
We introduce a population dynamics model, where individual genomes are represented by bit-strings. Selection is described by death probabilities which depend on these genomes, and new individuals continuously replace the ones that die,…
Compared to a neutral model, purifying selection distorts the structure of genealogies and hence alters the patterns of sampled genetic variation. Although these distortions may be common in nature, our understanding of how we expect…
High-throughput shotgun sequence data makes it possible in principle to accurately estimate population genetic parameters without confounding by SNP ascertainment bias. One such statistic of interest is the proportion of heterozygous sites…
Large populations may contain numerous simultaneously segregating polymorphisms subject to natural selection. Since selection acts on individuals whose fitness depends on many loci, different loci affect each other's dynamics. This leads to…
De novo assembly is the process of reconstructing the genome sequence of an organism from sequencing reads. Genome sequences are essential to biology, and assembly has been a central problem in bioinformatics for four decades. Until…
We investigate a continuous time, probability measure-valued dynamical system that describes the process of mutation-selection balance in a context where the population is infinite, there may be infinitely many loci, and there are weak…
When biological populations expand into new territory, the evolutionary outcomes can be strongly influenced by genetic drift, the random fluctuations in allele frequencies. Meanwhile, spatial variability in the environment can also…
Weak purifying selection, acting on many linked mutations, may play a major role in shaping patterns of molecular evolution in natural populations. Yet efforts to infer these effects from DNA sequence data are limited by our incomplete…
We study a simple solvable model describing the genesis of monomer sequences for hetero-polymers (such as proteins), as the result of the equilibration of a slow stochastic genetic selection process which is assumed to be driven by the…
We introduce a low dimensional function of the site frequency spectrum that is tailor-made for distinguishing coalescent models with multiple mergers from Kingman coalescent models with population growth, and use this function to construct…
The coding space of protein sequences is shaped by evolutionary constraints set by requirements of function and stability. We show that the coding space of a given protein family--the total number of sequences in that family--can be…
Increasing evidence suggests that chromosome folding and genetic expression are intimately connected. For example, the co-expression of a large number of genes can benefit from their spatial co-localization in the cellular space.…
We studied a relations between the triplet frequency composition of mitochondria genomes, and the phylogeny of their bearers. First, the clusters in 63dimensional space were developed due to $K$-means. Second, the clade composition of those…
The unwelcome evolution of malignancy during cancer progression emerges through a selection process in a complex heterogeneous population structure. In the present work, we investigate evolutionary dynamics in a phenotypically heterogeneous…
Computer simulations of complex population genetic models are an essential tool for making sense of the large-scale datasets of multiple genome sequences from a single species that are becoming increasingly available. A widely used approach…
The rules that specify how the information contained in DNA codes amino acids, is called "the genetic code". Using a simplified version of the Penna nodel, we are using computer simulations to investigate the importance of the genetic code…
The mean length and the variability of coding sequences for 48 genomes of bacteria and archaea were analyzed. It was found that the plotted data can be described by an angular area. This suggests the followings: a) The variability of a…
The amount of non-unique sequence (non-singletons) in a genome directly affects the difficulty of read alignment to a reference assembly for high throughput-sequencing data. Although a greater length increases the chance for reads being…