Related papers: Google matrix analysis of DNA sequences
We study the properties of random graphs where for each vertex a {\it neighbourhood} has been previously defined. The probability of an edge joining two vertices depends on whether the vertices are neighbours or not, as happens in Small…
We study networks constructed from gene expression data obtained from many types of cancers. The networks are constructed by connecting vertices that belong to each others' list of K-nearest-neighbors, with K being an a priori selected…
We consider Markov chains with random transition probabilities which, moreover, fluctuate randomly with time. We describe such a system by a product of stochastic matrices, $U(t)=M_t\cdots M_1$, with the factors $M_i$ drawn independently…
The frequency distributions of DNA k-mers are shaped by fundamental biological processes and offer a window into genome structure and evolution. Inspired by analogies to natural language, prior studies have attempted to model genomic k-mer…
We study scale free simple graphs with an exponent of the degree distribution $\gamma$ less than two. Generically one expects such extremely skewed networks -- which occur very frequently in systems of virtually or logically connected units…
This work proposes a markovian memoryless model for the DNA that simplifies enormously the complexity of it. We encode nucleotide sequences into symbolic sequences, called words, from which we establish meaningful length of words and group…
Most of the real world complex networks such as the Internet, World Wide Web and collaboration networks are huge; and to infer their structure and dynamics one requires handling large connectivity (adjacency) matrices. Also, to find out the…
We consider the problem of selecting important nodes in a random network, where the nodes connect to each other randomly with certain transition probabilities. The node importance is characterized by the stationary probabilities of the…
Interpreting natural language is an increasingly important task in computer algorithms due to the growing availability of unstructured textual data. Natural Language Processing (NLP) applications rely on semantic networks for structured…
The Internet is a complex network of interconnected routers and the existence of collective behavior such as congestion suggests that the correlations between different connections play a crucial role. It is thus critical to measure and…
Many real-world networks are intrinsically directed. Such networks include activation of genes, hyperlinks on the internet, and the network of followers on Twitter among many others. The challenge, however, is to create a network model that…
Distributions of triplets in some genetic sequences are examined and found to be well described by a 2-parameter Markov process with a sparse transition matrix. The variances of all the relevant parameters are not large, indicating that…
Complex networks have been studied extensively due to their relevance to many real systems as diverse as the World-Wide-Web (WWW), the Internet, energy landscapes, biological and social networks…
This paper proposed a new method to build the large scale DNA sequences search system based on web search engine technology. We give a very brief introduction for the methods used in search engine first. Then how to build a DNA search…
With the increasing number of texts made available on the Internet, many applications have relied on text mining tools to tackle a diversity of problems. A relevant model to represent texts is the so-called word adjacency (co-occurrence)…
Here, we propose a class of scale-free networks $G(t;m)$ with some intriguing properties, which can not be simultaneously held by all the theoretical models with power-law degree distribution in the existing literature, including (i)…
Many complex systems--from social and communication networks to biological networks and the Internet--are thought to exhibit scale-free structure. However, prevailing explanations rely on the constant addition of new nodes, an assumption…
One of the most influential recent results in network analysis is that many natural networks exhibit a power-law or log-normal degree distribution. This has inspired numerous generative models that match this property. However, more recent…
Word matches are often used in sequence comparison methods, either as a measure of sequence similarity or in the first search steps of algorithms such as BLAST or BLAT. The D2 statistic is the number of matches of words of k letters between…
We study a new class of matrix models, formulated on a lattice. On each site are $N$ states with random energies governed by a Gaussian random matrix Hamiltonian. The states on different sites are coupled randomly. We calculate the density…