English
Related papers

Related papers: Comparative statistical analysis of bacteria genom…

200 papers

The origin of long-range letter correlations in natural texts is studied using random walk analysis and Jensen-Shannon divergence. It is concluded that they result from slow variations in letter frequency distribution, which are a…

Computation and Language · Computer Science 2016-11-27 Dmitrii Y. Manin

The distributed genome hypothesis states that the set of genes in a population of bacteria is distributed over all individuals that belong to the specific taxon. It implies that certain genes can be gained and lost from generation to…

Probability · Mathematics 2010-11-08 F. Baumdicker , W. R. Hess , P. Pfaffelhuber

Since the completion of the human genome sequencing project in 2001, significant progress has been made in areas such as gene regulation editing and protein structure prediction. However, given the vast amount of genomic data, the segments…

Other Quantitative Biology · Quantitative Biology 2025-01-29 Wang Liang

Given two sequences over a finite alphabet $\mathcal{L}$, the $D_2$ statistic is the number of $m$-letter word matches between the two sequences. This statistic is used in bioinformatics for expressed sequence tag database searches. Here we…

Probability · Mathematics 2009-09-29 Conrad J. Burden , Miriam R. Kantorovitz , Susan R. Wilson

Lexicon based sentiment analysis usually relies on the identification of various words to which a numerical value corresponding to sentiment can be assigned. In principle, classifiers can be obtained from these algorithms by comparison with…

Computation and Language · Computer Science 2019-06-21 Mateus Machado , Evandro Ruiz , Kuruvilla Joseph Abraham

The typical process for classifying and submitting a newly sequenced virus to the NCBI database involves two steps. First, a BLAST search is performed to determine likely family candidates. That is followed by checking the candidate…

Genomics · Quantitative Biology 2016-03-22 Troy Hernandez , Jie Yang

This thesis aims at the logical analysis of discrete processes, in particular of such generated by gene regulatory networks. States, transitions and operators from temporal logics are expressed in the language of Formal Concept Analysis. By…

Molecular Networks · Quantitative Biology 2012-04-11 Johannes Wollbold

We show that the laws of autocorrelations decay in texts are closely related to applicability limits of language models. Using distributional semantics we empirically demonstrate that autocorrelations of words in texts decay according to a…

Computation and Language · Computer Science 2023-05-12 Nikolay Mikhaylovskiy , Ilya Churilov

Metagenomics sequencing is routinely applied to quantify bacterial abundances in microbiome studies, where the bacterial composition is estimated based on the sequencing read counts. Due to limited sequencing depth and DNA dropouts, many…

Methodology · Statistics 2019-04-26 Yuanpei Cao , Anru Zhang , Hongzhe Li

In condensed matter physics, simplified descriptions are obtained by coarse-graining the features of a system at a certain characteristic length, defined as the typical length beyond which some properties are no longer correlated. From a…

Genomics · Quantitative Biology 2018-04-18 Ivan Junier , Paul Frémont , Olivier Rivoire

By analyzing the spacing of genes on chromosomes, we find that transcriptional and RNA-processing regulatory sequences outside coding regions leave footprints on the distribution of intergenic distances. Using analogies between genes on…

Genomics · Quantitative Biology 2008-03-11 Rutger Hermsen , Pieter Rein ten Wolde , Sarah Teichmann

A theory of systems with long-range correlations based on the consideration of binary N-step Markov chains is developed. In the model, the conditional probability that the i-th symbol in the chain equals zero (or unity) is a linear function…

Data Analysis, Statistics and Probability · Physics 2016-09-08 O. V. Usatenko , V. A. Yampol'skii , K. E. Kechedzhy , S. S. Mel'nyk

We study the adaptation of Link Grammar Parser to the biomedical sublanguage with a focus on domain terms not found in a general parser lexicon. Using two biomedical corpora, we implement and evaluate three approaches to addressing unknown…

Computation and Language · Computer Science 2007-05-23 Sampo Pyysalo , Tapio Salakoski , Sophie Aubin , Adeline Nazarenko

The paradigm of large language models in natural language processing (NLP) has also shown promise in modeling biological languages, including proteins, RNA, and DNA. Both the auto-regressive generation paradigm and evaluation metrics have…

Biomolecules · Quantitative Biology 2025-07-04 Ke Liu , Shuaike Shen , Hao Chen

This article describes the results of a systematic in-depth study of the criteria used for word sense disambiguation. Our study is based on 60 target words: 20 nouns, 20 adjectives and 20 verbs. Our results are not always in line with some…

Computation and Language · Computer Science 2007-05-23 Laurent Audibert

Consider a random word $X^n=(X_1,\ldots ,X_n)$ in an alphabet consisting of $4$ letters, with the letters viewed either as $A$, $U$, $G$ and $C$ (i.e., nucleotides in an RNA sequence) or $\alpha$, $\bar{\alpha}$, $\beta$ and $\bar{\beta}$…

Group Theory · Mathematics 2022-01-20 Siddhartha Gadgil , Manjunath Krishnapur

Written language is a complex communication signal capable of conveying information encoded in the form of ordered sequences of words. Beyond the local order ruled by grammar, semantic and thematic structures affect long-range patterns in…

Physics and Society · Physics 2010-05-17 Marcelo A. Montemurro , Damian Zanette

Knowing the precise format of a program's input is a necessary prerequisite for systematic testing. Given a program and a small set of sample inputs, we (1) track the data flow of inputs to aggregate input fragments that share the same data…

Programming Languages · Computer Science 2017-08-30 Matthias Höschele , Alexander Kampmann , Andreas Zeller

Unraveling the evolutionary forces shaping bacterial diversity can today be tackled using a growing amount of genomic data. While the genome of eukaryotes is highly stable, bacterial genomes from cells of the same species highly vary in…

Populations and Evolution · Quantitative Biology 2015-03-19 Franz Baumdicker , Peter Pfaffelhuber

The task of text segmentation may be undertaken at many levels in text analysis---paragraphs, sentences, words, or even letters. Here, we focus on a relatively fine scale of segmentation, hypothesizing it to be in accord with a stochastic…