Related papers: Multifractal characterisation of complete genomes
We define multideterminantal probability measures, a family of probability measures on $[k]^n$ where $[k]=\{1,2,\dots,k\}$, generalizing determinantal measures (which correspond to the case $k=2$). We give examples coming from the positive…
How can we identify causal genetic mechanisms that govern bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype return high accuracy scores. However, attempts to…
We discuss properties of random fractals by means of a set of numbers that characterize their universal properties. This set is the generalized singularity specturm that consists of the usual spectrum of mulitfractal dimensions and the…
A natural representation of random graphs is the random measure. The collection of product random measures, their transformations, and non-negative test functions forms a general representation of the collection of non-negative weighted…
We introduce and study multiple partition structures which are sequences of probability measures on families of Young diagrams subjected to a consistency condition. The multiple partition structures are generalizations of Kingman's…
We define the complexity of DNA sequences as the information content per nucleotide, calculated by means of some Lempel-Ziv data compression algorithm. It is possible to use the statistics of the complexity values of the functional regions…
Deep generative models (DGMs) have recently demonstrated remarkable success in capturing complex probability distributions over graphs. Although their excellent performance is attributed to powerful and scalable deep neural networks, it is,…
In condensed matter physics, simplified descriptions are obtained by coarse-graining the features of a system at a certain characteristic length, defined as the typical length beyond which some properties are no longer correlated. From a…
The reconstruction of phylogenies from DNA or protein sequences is a major task of computational evolutionary biology. Common phenomena, notably variations in mutation rates across genomes and incongruences between gene lineage histories,…
Biological cells replicate their genomes in a well-planned manner. The DNA replication program of an organism determines the timing at which different genomic regions are replicated, with fundamental consequences for cell homeostasis and…
We introduce a nonparametric way to estimate the global probability density function for a random persistence diagram. Precisely, a kernel density function centered at a given persistence diagram and a given bandwidth is constructed. Our…
The n-point statistics of singularity strength variables for multiplicative branching processes is calculated from an analytic expression of the corresponding multivariate generating function. The key ingredient is a branching generating…
Next-generation sequencing technologies now constitute a method of choice to measure gene expression. Data to analyze are read counts, commonly modeled using Negative Binomial distributions. A relevant issue associated with this…
The rank ordered distribution of the codon usage frequencies for 123 bacteriae is best fitted by a three parameters function that is the sum of a constant, an exponential and a linear term in the rank n. The parameters depend (two…
We have found an analytic expression for the multivariate generating function governing all n-point statistics of random multiplicative cascade processes. The variable appropriate for this generating function is the logarithm of the energy…
In [3], we have introduced a probability measure to study the power and exponential sums for a certain coding system. The distribution function of the probability measure gives explicit formulas for the power and exponential sums.…
We study the stochastic dynamics of sequences evolving by single site mutations, segmental duplications, deletions, and random insertions. These processes are relevant for the evolution of genomic DNA. They define a universality class of…
Bacterial genomes and large-scale computer software projects both consist of a large number of components (genes or software packages) connected via a network of mutual dependencies. Components can be easily added or removed from individual…
The familiar cascade measures are sequences of random positive measures obtained on $[0,1]$ via $b$-adic independent cascades. To generalize them, this paper allows the random weights invoked in the cascades to take real or complex values.…
A causal set is a partially ordered set on a countably infinite ground-set such that each element is above finitely many others. A natural extension of a causal set is an enumeration of its elements which respects the order. We bring…