Related papers: Mutual information for examining correlations in D…
Fractal dimension is widely adopted in spatial databases and data mining, among others as a measure of dataset skewness. State-of-the-art algorithms for estimating the fractal dimension exhibit linear runtime complexity whether based on…
DNA copy number and mRNA expression are widely used data types in cancer studies, which combined provide more insight than separately. Whereas in existing literature the form of the relationship between these two types of markers is fixed a…
A promising new method for measuring intramolecular distances in solution uses small-angle X-ray scattering interference between gold nanocrystal labels (Mathew-Fenn et al, Science, 322, 446 (2008)). When applied to double stranded DNA, it…
The structure of the large scale distribution of the galaxies have been widely studied since the publication of the first catalogs. Since large redshift samples are available, their analyses seem to show fractal correlations up to the…
In this paper, we present a new approach to interpret deep learning models. By coupling mutual information with network science, we explore how information flows through feedforward networks. We show that efficiently approximating mutual…
We examine the Detrended Fluctuation Analysis (DFA), which is a well-established method for the detection of long-range correlations in time series. We show that deviations from scaling that appear at small time scales become stronger in…
Mutual information (MI) is a fundamental measure of statistical dependence between two variables, yet accurate estimation from finite data remains notoriously difficult. No estimator is universally reliable, and common approaches fail in…
The edit distance under the DCJ model can be computed in linear time for genomes with equal content or with Indels. But it becomes NP-Hard in the presence of duplications, a problem largely unsolved especially when Indels are considered. In…
Given two distinct datasets, an important question is if they have arisen from the the same data generating function or alternatively how their data generating functions diverge from one another. In this paper, we introduce an approach for…
Recent advances in high-throughput genomics technologies have resulted in the sequencing of large numbers of (near) complete genomes. These genome sequences are being mined for important functional elements, such as genes. They are also…
Femtoscopy is a powerful tool that can be used to investigate the space-time dimensions of the region from which the particles are emitted. When applied to high energy collisions this method is sensitive not only to quantum statistics, but…
Persistent homology allows us to create topological summaries of complex data. In order to analyse these statistically, we need to choose a topological summary and a relevant metric space in which this topological summary exists. While…
A variety of genome-wide profiling techniques are available to probe complementary aspects of genome structure and function. Integrative analysis of heterogeneous data sources can reveal higher-level interactions that cannot be detected…
Correlation testing provides a quick method of discriminating amongst potential terms to include in a nuclear mass formula or functional and is a necessary tool for further nuclear mass models; however a firm mathematical foundation of the…
We consider the task of detecting regulatory elements in the human genome directly from raw DNA. Past work has focused on small snippets of DNA, making it difficult to model long-distance dependencies that arise from DNA's 3-dimensional…
Labeling of DNA molecules is a fundamental technique for DNA visualization and analysis. This process was mathematically modeled in [1], where the received sequence indicates the positions of the used labels. In this work, we develop error…
We propose a statistical method to test whether two phylogenetic trees with given alignments are significantly incongruent. Our method compares the two distributions of phylogenetic trees given by the input alignments, instead of comparing…
Renormalization is an essential technique in field-theoretic descriptions of natural phenomena, where the absence of a UV-complete description yields an abundance of divergent quantities. While the renormalization prescription has been…
Existing sequence alignment algorithms use heuristic scoring schemes which cannot be used as objective distance metrics. Therefore one relies on measures like the p- or log-det distances, or makes explicit, and often simplistic, assumptions…
Mutual Information (MI) is a powerful statistical measure that quantifies shared information between random variables, particularly valuable in high-dimensional data analysis across fields like genomics, natural language processing, and…