Related papers: Time series model based on global structure of com…
A time series model of CDS sequences in complete genome is proposed. A map of DNA sequence to integer sequence is given. The correlation dimensions and Hurst exponents of CDS sequences in complete genome of bacteria are calculated. Using…
This paper considers three kinds of length sequences of the complete genome. Detrended fluctuation analysis, spectral analysis, and the mean distance spanned within time $L$ are used to discuss the correlation property of these sequences.…
In this paper we examine a number of methods for probing and understanding the large-scale structure of networks that evolve over time. We focus in particular on citation networks, networks of references between documents such as papers,…
Coding information is the main source of heterogeneity (non-randomness) in the sequences of bacterial genomes. This information can be naturally modeled by analysing cluster structures in the "in-phase" triplet distributions of relatively…
A time-dependent global fiber-bundle model of fracture with continuous damage is formulated in terms of a set of coupled non-linear differential equations. A first integral of this set is analytically obtained. The time evolution of the…
We consider an age-size structured cell population model based on the cell cycle length. The model is described by a first order partial differential equation with initial-boundary conditions. Using the theory of semigroups of positive…
The periodic transference of nucleotide strings in bacterial and archaeal complete genomes is investigated by using the metric representation and the recurrence plot method. The generated periodic correlation structures exhibit four kinds…
Protein distributions measured under a broad set of conditions in bacteria and yeast were shown to exhibit a common skewed shape, with variances depending quadratically on means. For bacteria these properties were reproduced by temporal…
A structural genetic model incorporating a modern understanding of the genome and common practice in genome-wide association studies is derived mathematically. The model shows the Haldane map distance as a direct consequence of the…
This is a review of a set of recent papers with some new data added. After a brief biological introduction a visualization scheme of the string composition of long DNA sequences, in particular, of bacterial complete genomes, will be…
The modeling of time series is becoming increasingly critical in a wide variety of applications. Overall, data evolves by following different patterns, which are generally caused by different user behaviors. Given a time series, we define…
Research into time series classification has tended to focus on the case of series of uniform length. However, it is common for real-world time series data to have unequal lengths. Differing time series lengths may arise from a number of…
The coding and noncoding length sequences constructed from a complete genome are characterised by multifractal analysis. The dimension spectrum $D_{q}$ and its derivative, the 'analogous' specific heat $C_{q}$, are calculated for the coding…
This work presents an introduction to feature-based time-series analysis. The time series as a data type is first described, along with an overview of the interdisciplinary time-series analysis literature. I then summarize the range of…
In condensed matter physics, simplified descriptions are obtained by coarse-graining the features of a system at a certain characteristic length, defined as the typical length beyond which some properties are no longer correlated. From a…
Mechanisms of immunity, and of the host-pathogen interactions in general are among the most fundamental problems of medicine, ecology, and evolution studies. Here, we present a microscopic, protein-level, sequence-based model of immune…
This paper describes characteristic features of networks reconstructed from gene expression time series data. Several null models are considered in order to discriminate between informations embedded in the network that are related to real…
In economics and many other forecasting domains, the real world problems are too complex for a single model that assumes a specific data generation process. The forecasting performance of different methods changes depending on the nature of…
The rapid development of high-throughput sequencing technologies has led to an explosive increase in biological sequence data, making sequence clustering a fundamental task in large-scale bioinformatics analyses. Unlike traditional…
This talk will review a little over a decade's research on applying certain stochastic models to biological sequence analysis. The models themselves have a longer history, going back over 30 years, although many novel variants have arisen…