English
Related papers

Related papers: Comparative statistical analysis of bacteria genom…

200 papers

Dependency distance minimization (DDm) is a word order principle favouring the placement of syntactically related words close to each other in sentences. Massive evidence of the principle has been reported for more than a decade with the…

Computation and Language · Computer Science 2021-02-02 Ramon Ferrer-i-Cancho , Carlos Gómez-Rodríguez

We present a statistical model of bacterial evolution based on the coupling between codon usage and tRNA abundance. Such a model interprets this aspect of the evolutionary process as a balance between the codon homogenization effect due to…

Statistical Mechanics · Physics 2007-05-23 Franco Bagnoli , Pietro Lio'

We perform differential expression analysis of high-throughput sequencing count data under a Bayesian nonparametric framework, removing sophisticated ad-hoc pre-processing steps commonly required in existing algorithms. We propose to use…

Applications · Statistics 2017-05-04 Siamak Zamani Dadaneh , Xiaoning Qian , Mingyuan Zhou

It is now a common practice to compare models of human language processing by predicting participant reactions (such as reading times) to corpora consisting of rich naturalistic linguistic materials. However, many of the corpora used in…

Computation and Language · Computer Science 2017-08-22 Richard Futrell , Edward Gibson , Hal Tily , Idan Blank , Anastasia Vishnevetsky , Steven T. Piantadosi , Evelina Fedorenko

The main statistical distributions applicable to the analysis of genome architecture and genome tracks are briefly discussed and critically assessed. Although the observed features in distributions of element lengths can be equally well…

Other Quantitative Biology · Quantitative Biology 2015-06-17 V. R. Chechetkin

The relationship between written and spoken words is convoluted in languages with a deep orthography such as English and therefore it is difficult to devise explicit rules for generating the pronunciations for unseen words. Pronunciation by…

Computation and Language · Computer Science 2011-09-22 Janne V. Kujala , Aleksi Keurulainen

We use an information-theoretic measure of linguistic similarity to investigate the organization and evolution of scientific fields. An analysis of almost 20M papers from the past three decades reveals that the linguistic similarity is…

Digital Libraries · Computer Science 2018-01-30 Laercio Dias , Martin Gerlach , Joachim Scharloth , Eduardo G. Altmann

We perform an exhaustive analysis of genome statistics for organisms, particularly extremophiles, growing in a wide range of physicochemical conditions. Specifically, we demonstrate how the correlation between the frequency of amino acids…

Genomics · Quantitative Biology 2013-09-19 Benjamin Greenbaum , Pradeep Kumar , Albert Libchaber

Learning language of protein sequences, which captures non-local interactions between amino acids close in the spatial structure, is a long-standing bioinformatics challenge, which requires at least context-free grammars. However, complex…

Formal Languages and Automata Theory · Computer Science 2019-03-20 Witold Dyrka , François Coste , Juliette Talibart

Autoregressive language models (LMs) map token sequences to probabilities. The usual practice for computing the probability of any character string (e.g. English sentences) is to first transform it into a sequence of tokens that is scored…

Computation and Language · Computer Science 2023-07-03 Nadezhda Chirkova , Germán Kruszewski , Jos Rozen , Marc Dymetman

Linguistic laws constitute one of the quantitative cornerstones of modern cognitive sciences and have been routinely investigated in written corpora, or in the equivalent transcription of oral corpora. This means that inferences of…

Physics and Society · Physics 2016-10-11 Ivan Gonzalez Torre , Bartolo Luque , Lucas Lacasa , Jordi Luque , Antoni Hernandez-Fernandez

This paper presents a new method for automatically detecting words with lexical gender in large-scale language datasets. Currently, the evaluation of gender bias in natural language processing relies on manually compiled lexicons of…

Computation and Language · Computer Science 2022-06-29 Marion Bartl , Susan Leavy

Recently long range correlations were detected in nucleotide sequences and in human writings by several authors. We undertake here a systematic investigation of two books, Moby Dick by H. Melville and Grimm's tales, with respect to the…

chao-dyn · Physics 2009-10-22 Werner Ebeling , Thorsten Pöschel

Newberry et al. (Detecting evolutionary forces in language change, Nature 551, 2017) tackle an important but difficult problem in linguistics, the testing of selective theories of language change against a null model of drift. Having…

Computation and Language · Computer Science 2020-05-08 Andres Karjus , Richard A. Blythe , Simon Kirby , Kenny Smith

String barcoding is a recently introduced technique for genomic-based identification of microorganisms. In this paper we describe the engineering of highly scalable algorithms for robust string barcoding. Our methods enable distinguisher…

Data Structures and Algorithms · Computer Science 2016-08-31 Bhaskar DasGupta , Kishori M. Konwar , Ion I. Mandoiu , Alex A. Shvartsman

The genome sequence contains the blueprint for governing cellular processes. While the availability of genomes has vastly increased over the last decades, experimental annotation of the various functional, non-coding and regulatory elements…

Genomics · Quantitative Biology 2024-04-10 Frederikke Isa Marin , Felix Teufel , Marc Horlacher , Dennis Madsen , Dennis Pultz , Ole Winther , Wouter Boomsma

The coding and noncoding length sequences constructed from a complete genome are characterised by multifractal analysis. The dimension spectrum $D_{q}$ and its derivative, the 'analogous' specific heat $C_{q}$, are calculated for the coding…

Biological Physics · Physics 2009-11-07 Zu-Guo Yu , Vo Anh , Ka-Sing Lau

While most current high-throughput DNA sequencing technologies generate short reads with low error rates, emerging sequencing technologies generate long reads with high error rates. A basic question of interest is the tradeoff between read…

Information Theory · Computer Science 2015-01-27 Ilan Shomorony , Thomas Courtade , David Tse

Auto-regulation, a process wherein a protein negatively regulates its own production, is a common motif in gene expression networks. Negative feedback in gene expression plays a critical role in buffering intracellular fluctuations in…

Subcellular Processes · Quantitative Biology 2014-05-16 Mohammad Soltani , Cesar Vargas , Niraj Kumar , Rahul Kulkarni , Abhyudai Singh

We propose an iterative algorithm to investigate the cooperative evolution dominated by information encoded within state spaces in a random quantum cellular automaton. Inspired by the 2-gram model in statistical linguistics, the updates of…

Quantum Physics · Physics 2025-04-22 Guanhua Chen , Yao Yao
‹ Prev 1 8 9 10 Next ›