Related papers: Sequence Heterogeneity Accelerates Protein Search …
Different numerical mappings of the DNA sequences have been studied using a new cluster-scaling method and the well known spectral methods. It is shown, in particular, that the nucleotide sequences in DNA molecules have robust…
It is not known how a cell manages to find a specific DNA sequence sufficiently fast to repair a broken chromosome through homologous recombination. I propose that the solution is based on a parallelized search implemented by freely…
We present a sequence-based probabilistic formalism that directly addresses co-operative effects in networks of interacting positions in proteins, providing significantly improved contact prediction, as well as accurate quantitative…
This paper describes a method to efficiently retrieve protein database sequences similar to a query sequence, while allowing for significant numbers of mutations. We call this method SEQR for SEQuence Retrieval. This approach increases the…
High-throughput genetic and epigenetic data are often screened for associations with an observed phenotype. For example, one may wish to test hundreds of thousands of genetic variants, or DNA methylation sites, for an association with…
Genetic recombination can produce heterogeneous phylogenetic histories within a set of homologous genes. Delineating recombination events is important in the study of molecular evolution, as inference of such events provides a clearer…
We wish to suggest an algorithm for biological search including DNA search. Our argument supposes that biological search be performed by quantum search.If we assume this, we can naturally answer the following long lasting puzzles such that…
At the core of high throughput DNA sequencing platforms lies a bio-physical surface process that results in a random geometry of clusters of homogenous short DNA fragments typically hundreds of base pairs long - bridge amplification. The…
DNA sequence encoding is fundamental to gene function prediction, protein synthesis, and diverse downstream biological tasks. Despite the substantial progress achieved by large-scale DNA sequence pretraining, existing studies have…
Genome analysis fundamentally starts with a process known as read mapping, where sequenced fragments of an organism's genome are compared against a reference genome. Read mapping is currently a major bottleneck in the entire genome analysis…
How DNA-binding proteins locate specific genomic targets remains a central challenge in molecular biology. Traditional protein-centric approaches, which rely on wet-lab experiments and visualization techniques, often lack genome-wide…
The model of facilitated diffusion describes how DNA-binding proteins, such as transcription factors (TFs), find their chromosomal targets by combining 3D diffusion through the cytoplasm and 1D sliding along nonspecific DNA sequences. The…
The interaction between proteins and DNA is a key driving force in a significant number of biological processes such as transcriptional regulation, repair, recombination, splicing, and DNA modification. The identification of DNA-binding…
We study a model for a protein searching for a target, using facilitated diffusion, on a DNA molecule confined in a finite volume. The model includes three distinct pathways for facilitated diffusion: (a) sliding - in which the protein…
Mapping between sequence and structure is currently an open problem in structural biology. Despite many experimental and computational efforts it is not clear yet how the structure is encoded in the sequence. Answering this question may…
The precision of biochemical signaling is limited by randomness in the diffusive arrival of molecules at their targets. For proteins binding to the specific sites on the DNA and regulating transcription, the ability of the proteins to…
Much of the on-going statistical analysis of DNA sequences is focused on the estimation of characteristics of coding and non-coding regions that would possibly allow discrimination of these regions. In the current approach, we concentrate…
The mechanical model based on beads and springs, which we recently proposed to study non-specific DNA-protein interactions [J. Chem. Phys. 130, 015103 (2009)], was improved by describing proteins as sets of interconnected beads instead of…
We provide an overview of current approaches to DNA-based storage system design and accompanying synthesis, sequencing and editing methods. We also introduce and analyze a suite of new constrained coding schemes for both archival and random…
The sequence of a protein is not only constrained by its physical and biochemical properties under current selection, but also by features of its past evolutionary history. Understanding the extent and the form that these evolutionary…