English
Related papers

Related papers: A Note on Zipf's Law, Natural Languages, and Nonco…

200 papers

Ensembl's human non-coding and protein coding genes are used to automatically find DNA pattern motifs. The Backus-Naur form (BNF) grammar for regular expressions (RE) is used by genetic programming to ensure the generated strings are legal.…

Biomolecules · Quantitative Biology 2010-02-02 W. B. Langdon , Olivia Sanchez Graillet , A. P. Harrison

Language models, especially transformer-based ones, have achieved colossal success in NLP. To be precise, studies like BERT for NLU and works like GPT-3 for NLG are very important. If we consider DNA sequences as a text written with an…

Genomics · Quantitative Biology 2026-01-21 Musa Nuri Ihtiyar , Arzucan Ozgur

Natural code is known to be very repetitive (much more so than natural language corpora); furthermore, this repetitiveness persists, even after accounting for the simpler syntax of code. However, programming languages are very expressive,…

Computation and Language · Computer Science 2019-10-10 Casey Casalnuovo , Kevin Lee , Hulin Wang , Prem Devanbu , Emily Morgan

The DNA storage channel is considered, in which a codeword is comprised of $M$ unordered DNA molecules. At reading time, $N$ molecules are sampled with replacement, and then each molecule is sequenced. A coded-index concatenated-coding…

Information Theory · Computer Science 2022-05-23 Nir Weinberger

We use the formulation of equilibrium statistical mechanics in order to study some important characteristics of language. Using a simple expression for the Hamiltonian of a language system, which is directly implied by the Zipf law, we are…

Physics and Society · Physics 2009-11-11 Kosmas Kosmidis , Alkiviadis Kalampokis , Panos Argyrakis

The genetic code is the function from the set of codons to the set of amino acids by which a DNA sequence encodes proteins. Since the codons also influence the shape of the DNA molecule itself, the same sequence that encodes a protein also…

Other Quantitative Biology · Quantitative Biology 2020-03-04 Alex Kasman , Brenton LeMesurier

Despite being a paradigm of quantitative linguistics, Zipf's law for words suffers from three main problems: its formulation is ambiguous, its validity has not been tested rigorously from a statistical point of view, and it has not been…

Applications · Statistics 2016-02-17 Isabel Moreno-Sánchez , Francesc Font-Clos , Álvaro Corral

We propose the Transcendental Encoding Conjecture for decision problems, which asserts that every language in complexity class P encodes to an algebraic real (possibly rational or algebraic irrational) under its binary characteristic…

Computational Complexity · Computer Science 2025-06-26 Anand Kumar Keshavan , Sunu Engineer

The formation of DNA loops by proteins and protein complexes is ubiquitous to many fundamental cellular processes, including transcription, recombination, and replication. Here we review recent advances in understanding the properties of…

Biomolecules · Quantitative Biology 2007-05-23 Leonor Saiz , Jose M. G. Vilar

In our paper selected linguistic features of genomes to study the statistics of the gene codes are considered. We present the information theory from which it follows that if the system is described by distributions of hyperbolic type it…

Genomics · Quantitative Biology 2014-07-10 Krystyna Lukierska-Walasek , Krzysztof Topolski , Krzysztof Trojanowski

In this article, we evaluate computational models of natural language with respect to the universal statistical behaviors of natural language. Statistical mechanical analyses have revealed that natural language text is characterized by…

Computation and Language · Computer Science 2019-06-25 Shuntaro Takahashi , Kumiko Tanaka-Ishii

Signed languages are the primary means of communication for many deaf and hard of hearing individuals. Since signed languages exhibit all the fundamental linguistic properties of natural language, we believe that tools and theories of…

Computation and Language · Computer Science 2021-07-26 Kayo Yin , Amit Moryossef , Julie Hochgesang , Yoav Goldberg , Malihe Alikhani

We present a simple structure based model of how words are formed from morphemes. The model explains two major empirical facts: the typical distribution of word lengths and the appearance of Zipf like rank frequency curves. In contrast to…

Methodology · Statistics 2025-12-16 Vladimir Berman

Language, which allows complex ideas to be communicated through symbolic sequences, is a characteristic feature of our species and manifested in a multitude of forms. Using large written corpora for many different languages and scripts, we…

Computation and Language · Computer Science 2018-01-17 Md Izhar Ashraf , Sitabhra Sinha

Evolution consists of distinct stages: cosmological, biological, linguistic. Since biology verges on natural sciences and linguistics, we expect that it shares structures and features from both forms of knowledge. Indeed, in DNA we…

Other Quantitative Biology · Quantitative Biology 2019-10-01 Argyris Nicolaidis , Fotis Psomopoulos

Shannon information (SI) and its special case, divergence, are defined for a DNA sequence in terms of probabilities of chemical words in the sequence and are computed for a set of complete genomes highly diverse in length and composition.…

Genomics · Quantitative Biology 2009-11-10 Hong-Da Chen , Chang-Heng Chang , Li-Ching Hsieh , Hoong-Chien Lee

The availability of large datasets requires an improved view on statistical laws in complex systems, such as Zipf's law of word frequencies, the Gutenberg-Richter law of earthquake magnitudes, or scale-free degree distribution in networks.…

Data Analysis, Statistics and Probability · Physics 2019-04-30 Martin Gerlach , Eduardo G. Altmann

Tokenization is a crucial step in processing protein sequences for machine learning models, as proteins are complex sequences of amino acids that require meaningful segmentation to capture their functional and structural properties.…

Computation and Language · Computer Science 2024-11-27 Burak Suyunu , Enes Taylan , Arzucan Özgür

Researchers have observed that the frequencies of leading digits in many man-made and naturally occurring datasets follow a logarithmic curve, with digits that start with the number 1 accounting for $\sim 30\%$ of all numbers in the dataset…

Computation and Language · Computer Science 2022-12-22 Leo Hsu , Visar Berisha

Some authors have recently argued that a finite-size scaling law for the text-length dependence of word-frequency distributions cannot be conceptually valid. Here we give solid quantitative evidence for the validity of such scaling law,…

Data Analysis, Statistics and Probability · Physics 2018-04-12 Alvaro Corral , Francesc Font-Clos
‹ Prev 1 8 9 10 Next ›