Related papers: Protein-DNA computation by stochastic assembly cas…
We introduce a differentiable clustering method based on stochastic perturbations of minimum-weight spanning forests. This allows us to include clustering in end-to-end trainable pipelines, with efficient gradients. We show that our method…
Principal component analysis (PCA) aims at estimating the direction of maximal variability of a high-dimensional dataset. A natural question is: does this task become easier, and estimation more accurate, when we exploit additional…
Complex DNA topological structures, including polymer loops, are frequently observed in biological processes when protein molecules simultaneously bind to several distant sites on DNA. However, the molecular mechanisms of formation of these…
Genetic information is stored in a linear sequence of base-pairs; however, thermal fluctuations and complex DNA conformations such as folds and loops make it challenging to order genomic material for in vitro analysis. In this work, we…
By means of the concept of Factorial Moments we examine DNA sequences from Yeast to distinguish coding and non-coding regions. It is found that the FM may be a powerful tool for analysis of DNA sequences. PACS numbers:…
A computational approach by an implementation of the Principle Component Analysis (PCA) with K-means and Gaussian Mixture (GM) clustering methods from Machine Learning (ML) algorithms to identify structural and dynamical heterogeneities of…
Ribonucleic acid (RNA) is involved in many regulatory and catalytic processes in the cell. The function of any RNA molecule is intimately related with its structure. In-line probing experiments provide valuable structural datasets for a…
It is a standard exercise in mechanical engineering to infer the external forces and torques on a body from its static shape and known elastic properties. Here we apply this kind of analysis to distorted double-helical DNA in complexes with…
DNA nanotechnology uses predictable interactions of nucleic acids to precisely engineer complex nanostructures. Characterizing these self-assembled structures at the single-structure level is crucial for validating their design and…
The log-det distance between two aligned DNA sequences was introduced as a tool for statistically consistent inference of a gene tree under simple non-mixture models of sequence evolution. Here we prove that the log-det distance, coupled…
Fourier PCA is Principal Component Analysis of a matrix obtained from higher order derivatives of the logarithm of the Fourier transform of a distribution.We make this method algorithmic by developing a tensor decomposition method for a…
Stochastic simulation can make the molecular processes of cellular control more vivid than the traditional differential-equation approach by generating typical system histories instead of just statistical measures such as the mean and…
As a possible implementation of data storage using DNA, multiple strands of DNA are stored in a liquid container so that, in the future, they can be read by an array of DNA readers in parallel. These readers will sample the strands with…
Mode separation, namely how sharply a distribution fragments into barrier-separated clusters, is a fundamental geometric property of densities, difficult to quantify in high dimensions. It is structurally distinct from dispersion, yet…
DNA is an astonishing material that can be used as a molecular building block to construct periodic arrays and devices with nanoscale accuracy and precision. Here, we present simple bead-spring model of DNA nanostars having three, four and…
The field of complex self-assembly is moving toward the design of multi-particle structures consisting of thousands of distinct building blocks. To exploit the potential benefits of structures with such `addressable complexity,' we need to…
By using the Jensen-Shannon divergence, genomic DNA can be divided into compositionally distinct domains through a standard recursive segmentation procedure. Each domain, while significantly different from its neighbours, may however share…
In this paper, we propose and study a Nystr\"om based approach to efficient large scale kernel principal component analysis (PCA). The latter is a natural nonlinear extension of classical PCA based on considering a nonlinear feature map or…
Owing to its longevity and enormous information density, DNA, the molecule encoding biological information, has emerged as a promising archival storage medium. However, due to technological constraints, data can only be written onto many…
The adsorption or adhesion of large particles (proteins, colloids, cells, >...) at the liquid-solid interface plays an important role in many diverse applications. Despite the apparent complexity of the process, two features are…