Related papers: Statistical analysis of simple repeats in the huma…
This paper presents a novel method to segment/decode DNA sequences based on n-grams statistical language model. Firstly, we find the length of most DNA 'words' is 12 to 15 bps by analyzing the genomes of 12 model species. Then we design an…
In this paper, we present a novel design strategy of DNA codes with length $3n$ over the non-chain ring $R=\mathbb{Z}_4+u\mathbb{Z}_4+u^2\mathbb{Z}_4$ with $64$ elements and $u^3=1$, where $n$ denotes the length of a code over $R$. We first…
In this work we study reverse complementary genomic word pairs in the human DNA, by comparing both the distance distribution and the frequency of a word to those of its reverse complement. Several measures of dissimilarity between distance…
Summary: Human alpha satellite and satellite 2/3 contribute to several percent of the human genome. However, identifying these sequences with traditional algorithms is computationally intensive. Here we develop dna-brnn, a recurrent neural…
We investigate the localization property of an electron in the disordered two-chain system (ladder model) with long-range correlation as a simple model for electronic property in DNA sequence. The chains are constructed by repetition of the…
Genetic sequences are known to possess non-trivial composition together with symmetries in the frequencies of their components. Recently, it has been shown that symmetry and structure are hierarchically intertwined in DNA, suggesting a…
During evolution of microorganisms genomes underwork have different changes in their lengths, gene orders, and gene contents. Investigating these structural rearrangements helps to understand how genomes have been modified over time. Some…
We investigate the conformations of DNA-like stiff chains, characterized by contour length ($L$) and persistence length ($l_p$), in a variety of crowded environments containing mono disperse soft spherical (SS) and spherocylindrical (SC)…
High read depth can be used to assemble short sequence repeats. The existing genome assemblers fail in repetitive regions of longer than average read. I propose a new algorithm for a DNA assembly which uses the relative frequency of reads…
Repetitions within a given genealogical tree provides some information about the degree of consanguineity of a population. They can be analyzed with techniques usually employed in statistical physics when dealing with fixed point…
Gene sequences in the vicinity of splice sites are found to possess dinucleotide periodicities, especially RR and YY, with the period close to the pitch of nucleosome DNA. This confirms previously reported finding about preferential…
The human genome remains incomplete, with multi-megabase sized gaps representing the endogenous centromeres and other heterochromatic regions. These regions are commonly enriched with long arrays of near-identical tandem repeats, known as…
The dynamic spatial redistribution of individuals is a key driving force of various spatiotemporal phenomena on geographical scales. It can synchronise populations of interacting species, stabilise them, and diversify gene pools [1-3].…
A new family of compound Poisson distribution functions from statistical linguistic is used to study the n-tuples and nucleotide composition features of DNA sequences. The relative frequency distribution of the 6-tuples and 7- tuples…
While the behavior of double stranded DNA at mesoscopic scales is fairly well understood, less is known about its relation to the rich mechanical properties in the base-pair scale, which is crucial, for instance, to understand DNA-protein…
The human genotope is the convex hull of all allele frequency vectors that can be obtained from the genotypes present in the human population. In this paper we take a few initial steps towards a description of this object, which may be…
Most living systems rely on double-stranded DNA (dsDNA) to store their genetic information and perpetuate themselves. This biological information has been considered the main target of evolution. However, here we show that symmetries and…
Biologists have long sought a way to explain how statistical properties of genetic sequences emerged and are maintained through evolution. On the one hand, non-random structures at different scales indicate a complex genome organisation. On…
Advances in DNA nanotechnology have stimulated the search for simple motifs that can be used to control the properties of DNA nanostructures. One such motif, which has been used extensively in structures such as polyhedral cages,…
The high linear charge density of 20-base-pair oligomers of DNA is shown to lead to a striking non-monotonic dependence of the long-time self-diffusion on the concentration of the DNA in low-salt conditions. This generic non-monotonic…