Related papers: Construction of Multiple Constrained DNA Codes
In this review paper, we delve into the nascent field of molecular data storage, focusing on system implementations and code constructions. We start by providing an overview of basic concepts in synthetic and computational biology.…
The two strands of a DNA molecule with a repetitive sequence can pair into many different basepairing patterns. For perfectly periodic sequences, early bulk experiments of Poerschke indicate the existence of a sliding process, permitting…
Biological data mainly comprises of Deoxyribonucleic acid (DNA) and protein sequences. These are the biomolecules which are present in all cells of human beings. Due to the self-replicating property of DNA, it is a key constitute of genetic…
Two new constructions are presented for coils and snakes in the hypercube. Improvements are made on the best known results for snake-in-the-box coils of dimensions 9, 10 and 11, and for some other circuit codes of dimensions between 8 and…
The number of protein structures is far less than the number of sequences. By imposing simple generic features of proteins (low energy and compaction) on all possible sequences we show that the structure space is sparse compared to the…
The statistical mechanics of heteropolymer structure formation is studied in the context of RNA secondary structures. A designed RNA sequence biased energetically towards a particular native structure (a hairpin) is used to study the…
DNA-based data storage systems face practical challenges due to the high cost of DNA synthesis. A strategy to address the problem entails encoding data via topological modifications of the DNA sugar-phosphate backbone. The DNA Punchcards…
Motivation: Predicting the secondary structure of an RNA sequence is useful in many applications. Existing algorithms (based on dynamic programming) suffer from a major limitation: their runtimes scale cubically with the RNA length, and…
We modify and extend the recently developed statistical mechanical model for predicting the thermodynamic properties of chain molecules having noncovalent double-stranded conformations, as in RNA or ssDNA, and $\beta-$sheets in protein, by…
The coverage depth problem in DNA data storage is about minimizing the expected number of reads until all data is recovered. When they exist, MDS codes offer the best performance in this context. This paper focuses on the scenario where the…
Binary self-dual codes with large minimum distances, such as the extended Hamming code and the Golay code, are fascinating objects in the coding theory. They are closely related to sporadic simple groups, lattices and invariant theory. A…
The formation of DNA loops by proteins and protein complexes that bind at distal DNA sites plays a central role in many cellular processes, such as transcription, recombination, and replication. Here we review the basic thermodynamic…
This paper presents a novel method to segment/decode DNA sequences based on n-grams statistical language model. Firstly, we find the length of most DNA 'words' is 12 to 15 bps by analyzing the genomes of 12 model species. Then we design an…
Segmental duplications (SDs), or low-copy repeats (LCR), are segments of DNA greater than 1 Kbp with high sequence identity that are copied to other regions of the genome. SDs are among the most important sources of evolution, a common…
Protein structures in nature often exhibit a high degree of regularity (secondary structures, tertiary symmetries, etc.) absent in random compact conformations. We demonstrate in a simple lattice model of protein folding that structural…
In the human genomes, recombination frequency between homologous chromosomes during meiosis is highly correlated with their physical length while it differs significantly when their coding density is considered. Furthermore, it has been…
DNA supercoiling, the under or overwinding of DNA, is a key physical mechanism both participating to compaction of bacterial genomes and making genomic sequences adopt various structural forms. DNA supercoiling may lead to the formation of…
In this paper, fundamental limits in sequencing of a set of closely related DNA molecules are addressed. This problem is called pooled-DNA sequencing which encompasses many interesting problems such as haplotype phasing, metageomics, and…
Several processes in the cell, such as gene regulation, start when key proteins recognise and bind to short DNA sequences. However, as these sequences can be hundreds of million times shorter than the genome, they are hard to find by simple…
Our work is concerned with the generation and targeted design of RNA, a type of genetic macromolecule that can adopt complex structures which influence their cellular activities and functions. The design of large scale and complex…