Related papers: Experimental demonstration of kinetic proofreading…
The potential of a double nanopore system to determine DNA barcodes has been demonstrated experimentally. By carrying out Brownian dynamics simulation on a coarse-grained model DNA with protein tag (barcodes) at known locations along the…
We describe a strategy for constructing codes for DNA-based information storage by serial composition of weighted finite-state transducers. The resulting state machines can integrate correction of substitution errors; synchronization by…
We consider the hashing mechanism for constructing binary embeddings, that involves pseudo-random projections followed by nonlinear (sign function) mappings. The pseudo-random projection is described by a matrix, where not all entries are…
The two strands of a DNA molecule with a repetitive sequence can pair into many different basepairing patterns. For perfectly periodic sequences, early bulk experiments of Poerschke indicate the existence of a sliding process, permitting…
Labeling of DNA molecules is a fundamental technique for DNA visualization and analysis. This process was mathematically modeled in [1], where the received sequence indicates the positions of the used labels. In this work, we develop error…
DNA is emerging as an increasingly attractive medium for data storage due to a number of important and unique advantages it offers, most notably the unprecedented durability and density. While the technology is evolving rapidly, the…
Gene finding is the task of identifying the locations of coding sequences within the vast amount of genetic code contained in the genome. With an ever increasing quantity of raw genome sequences, gene finding is an important avenue towards…
In the postgenome era many efforts have been dedicated to systematically elucidate the complex web of interacting genes and proteins. These efforts include experimental and computational methods. Microarray technology offers an opportunity…
Lineage tracing, the determination and mapping of progeny arising from single cells, is an important approach enabling the elucidation of mechanisms underlying diverse biological processes ranging from development to disease. We developed a…
DNA is now firmly established as a versatile and robust platform for achieving synthetic nanostructures. While the folding of single molecules into complex structures is routinely achieved through engineering basepair sequences, much less…
One of the causes of high fidelity of copying in biological systems is kinetic discrimination. In this mechanism larger dissipation and copying velocity result in improved copying accuracy. We consider a model of a polymerase which…
Synthetic data has been proposed as a solution to address the issue of high-quality data scarcity in the training of large language models (LLMs). Studies have shown that synthetic data can effectively improve the performance of LLMs on…
In order to improve the accuracy of molecular dynamics simulations, classical force fields are supplemented with a kernel-based machine learning method trained on quantum-mechanical fragment energies. As an example application, a…
Protein structure prediction has been a grand challenge for over 50 years, owing to its broad scientific and application interests. There are two primary types of modeling algorithms, template-free modeling and template-based modeling. The…
Error backpropagation is a highly effective mechanism for learning high-quality hierarchical features in deep networks. Updating the features or weights in one layer, however, requires waiting for the propagation of error signals from…
The way to infer well-supported phylogenetic trees that precisely reflect the evolutionary process is a challenging task that completely depends on the way the related core genes have been found. In previous computational biology studies,…
Motivated by average-case trace reconstruction and coding for portable DNA-based storage systems, we initiate the study of \emph{coded trace reconstruction}, the design and analysis of high-rate efficiently encodable codes that can be…
State-of-the-art, high capacity deep neural networks not only require large amounts of labelled training data, they are also highly susceptible to label errors in this data, typically resulting in large efforts and costs and therefore…
In this work the computer modeling has been used to show that longer ligands allow biological cells (e.g., blood platelets) to withstand stronger flows after their adhesion to solid walls. Mechanistic model of polymer-mediated…
DNA stretching experiments are usually interpreted using the worm-like chain model; the persistence length A appearing in the model is then interpreted as the elastic stiffness of the double helix. In fact the persistence length obtained by…