Related papers: Spectral sum rules and search for periodicities in…
Stochastic neural networks are a prototypical computational device able to build a probabilistic representation of an ensemble of external stimuli. Building on the relationship between inference and learning, we derive a synaptic plasticity…
Statistical inference of genetic regulatory networks is essential for understanding temporal interactions of regulatory elements inside the cells. For inferences of large networks, identification of network structure is typical achieved…
We derive a formalism to describe the scattering of polarized radiation over the full spectral range encompassed by atomic transitions belonging to the same spectral series (e.g., the H I Lyman and Balmer series, the UV multiplets of Fe I…
It has been observed that an interesting class of non-Gaussian stationary processes is obtained when in the harmonics of a signal with random amplitudes and phases, frequencies can also vary randomly. In the resulting models, the…
The seriation problem seeks to reorder a set of elements given pairwise similarity information, so that elements with higher similarity are closer in the resulting sequence. When a global ordering consistent with the similarity information…
Much evolutionary information is stored in the fluctuations of protein length distributions. The genome size and non-coding DNA content can be calculated based only on the protein length distributions. So there is intrinsic relationship…
Properties of solutions of generic hyperbolic systems with multiple characteristics with diagonalizable principal part are investigated. Solutions are represented as a Picard series with terms in the form of iterated Fourier integral…
Various approaches to alignment-free sequence comparison are based on the length of exact or inexact word matches between two input sequences. Haubold {\em et al.} (2009) showed how the average number of substitutions between two DNA…
In this article, we review existing probabilistic models for modeling abundance of fixed-length strings (k-mers) in DNA sequencing data. These models capture dependence of the abundance on various phenomena, such as the size and repeat…
Sequential pattern discovery is a well-studied field in data mining. Episodes are sequential patterns describing events that often occur in the vicinity of each other. Episodes can impose restrictions to the order of the events, which makes…
Shannon information (SI) and its special case, divergence, are defined for a DNA sequence in terms of probabilities of chemical words in the sequence and are computed for a set of complete genomes highly diverse in length and composition.…
Frequent sequence mining methods often make use of constraints to control which subsequences should be mined. A variety of such subsequence constraints has been studied in the literature, including length, gap, span, regular-expression, and…
Fourier spectral estimates and, to a lesser extent, the autocorrelation function are the primary tools to detect periodicities in experimental data in the physical and biological sciences. We propose a new method which is more reliable than…
In solar physics, especially in exploratory stages of research, it is often necessary to compare the power spectra of two or more time series. One may, for instance, wish to estimate what the power spectrum of the combined data sets might…
With recent high-throughput technology we can synthesize large heterogeneous collections of DNA structures, and also read them all out precisely in a single procedure. Can we use these tools, not only to do things faster, but also to devise…
Biological sequences do not come at random. Instead, they appear with particular frequencies that reflect properties of the associated system or phenomenon. Knowing how biological sequences are distributed in sequence space is thus a…
A stretched DNA molecule which is also under- or overwound, undergoes a buckling transition forming intertwined looped domains called plectonemes. Here we develop a simple theory that extends the two-phase model of stretched supercoiled…
DNA-based storage has emerged as a promising alternative to traditional data storage methods, offering unmatched advantages in data density, longevity, and sustainability. Two main approaches have developed: in-vitro storage, where…
The identification of increasingly smaller signal from objects observed with a non-perfect instrument in a noisy environment poses a challenge for a statistically clean data analysis. We want to compute the probability of frequencies…
Shannon entropy is widely used to measure the complexity of DNA sequences but suffers from saturation effects that limit its discriminative power for long uniform segments. We introduce a novel metric, the entropy rank ratio R, which…