English
Related papers

Related papers: The Capacity of Some P\'olya String Models

200 papers

It is known that the majority of the human genome consists of repeated sequences. Furthermore, it is believed that a significant part of the rest of the genome also originated from repeated sequences and has mutated to its current form. In…

Information Theory · Computer Science 2014-01-21 Farzad Farnoud , Moshe Schwartz , Jehoshua Bruck

Motivated by DNA storage in living organisms, and by known biological mutation processes, we study the reverse-complement string-duplication system. We fully classify the conditions under which the system has full expressiveness, for all…

Information Theory · Computer Science 2021-12-23 Eyar Ben-Tolila , Moshe Schwartz

Genomic evolution can be viewed as string-editing processes driven by mutations. An understanding of the statistical properties resulting from these mutation processes is of value in a variety of tasks related to biological sequence data,…

Information Theory · Computer Science 2018-12-07 Hao Lou , Farzad Farnoud , Moshe Schwartz , Jehoshua Bruck

Modern language models define distributions over strings, but downstream tasks often require different output formats. For instance, a model that generates byte-pair strings does not directly produce word-level predictions, and a DNA model…

Computation and Language · Computer Science 2026-03-09 Vésteinn Snæbjarnarson , Samuel Kiegeland , Tianyu Liu , Reda Boumasmoud , Ryan Cotterell , Tim Vieira

This paper proposes a novel approach for statistical modelling of a continuous random variable $X$ on $[0, 1)$, based on its digit representation $X=.X_1X_2\ldots$. In general, $X$ can be coupled with a latent random variable $N$ so that…

Methodology · Statistics 2025-12-10 Mario Beraha , Jesper Møller

Information capacity of a symbol sequence is a measure of the unexpectedness of a continuation of given string of symbols. Continuation of a string is determined through the maximum entropy of the reconstructed frequency dictionary; the…

Genomics · Quantitative Biology 2007-05-23 Michael G. Sadovsky

The majority of the human genome consists of repeated sequences. An important type of repeated sequences common in the human genome are tandem repeats, where identical copies appear next to each other. For example, in the sequence…

Information Theory · Computer Science 2016-11-18 Siddharth Jain , Farzad Farnoud , Jehoshua Bruck

In this article, we review existing probabilistic models for modeling abundance of fixed-length strings (k-mers) in DNA sequencing data. These models capture dependence of the abundance on various phenomena, such as the size and repeat…

Quantitative Methods · Quantitative Biology 2022-01-03 Askar Gafurov , Tomáš Vinař , Broňa Brejová

The theme in this paper is a composition of random graphs and P\'olya urns. The random graphs are generated through a small structure called the seed. Via P\'olya urns, we study the asymptotic degree structure in a random $m$-ary hooking…

Probability · Mathematics 2024-08-07 Kiran R. Bhutani , Ravi Kalpathy , Hosam Mahmoud

In an attempt to explain the uniqueness of the coding mechanism of living cells as contrasted with multi-species structure of ecosystems we examine two models of individuals with some replicative properties. In the first model the system…

Statistical Mechanics · Physics 2009-10-31 A. Lipowski

We consider the problem of coding for the substring channel, in which information strings are observed only through their (multisets of) substrings. Due to existing DNA sequencing techniques and applications in DNA-based storage systems,…

Information Theory · Computer Science 2024-03-27 Yonatan Yehezkeally , Nikita Polyanskii

We introduce an extension of the P\'olya tree approach for constructing distributions on the space of probability measures. By using optional stopping and optional choice of splitting variables, the construction gives rise to random…

Statistics Theory · Mathematics 2010-10-05 Wing H. Wong , Li Ma

Genetic Algorithms are introduced as a search method for finding string vacua with viable phenomenological properties. It is shown, by testing them against a class of Free Fermionic models, that they are orders of magnitude more efficient…

High Energy Physics - Theory · Physics 2015-06-19 Steven Abel , John Rizos

In this work we give specific examples of competition models, with six and eight species, whose three-dimensional dynamics naturally leads to the formation of string networks with junctions, associated with regions that have a high…

Biological Physics · Physics 2017-02-17 P. P. Avelino , D. Bazeia , L. Losano , J. Menezes , B. F. de Oliveira

The relationship between the quality of a string, as judged by a human reader, and its probability, $p(\boldsymbol{y})$ under a language model undergirds the development of better language models. For example, many popular algorithms for…

Computation and Language · Computer Science 2024-10-29 Naaman Tan , Josef Valvoda , Tianyu Liu , Anej Svete , Yanxia Qin , Kan Min-Yen , Ryan Cotterell

Many practical modeling problems involve discrete data that are best represented as draws from multinomial or categorical distributions. For example, nucleotides in a DNA sequence, children's names in a given state and year, and text…

Machine Learning · Statistics 2015-06-22 Scott W. Linderman , Matthew J. Johnson , Ryan P. Adams

We investigate the performance of large language models on repetitive deterministic prediction tasks and study how the sequence accuracy rate scales with output length. Each such task involves repeating the same operation n times. Examples…

Artificial Intelligence · Computer Science 2025-11-25 Wanda Hou , Leon Zhou , Hong-Ye Hu , Yubei Chen , Yi-Zhuang You , Xiao-Liang Qi

The problem of reconstructing strings from their substring spectra has a long history and in its most simple incarnation asks for determining under which conditions the spectrum uniquely determines the string. We study the problem of coded…

Information Theory · Computer Science 2019-04-24 Ryan Gabrys , Olgica Milenkovic

What can large language models learn? By definition, language models (LM) are distributions over strings. Therefore, an intuitive way of addressing the above question is to formalize it as a matter of learnability of classes of…

Computation and Language · Computer Science 2025-01-14 Nadav Borenstein , Anej Svete , Robin Chan , Josef Valvoda , Franz Nowak , Isabelle Augenstein , Eleanor Chodroff , Ryan Cotterell

We show a possibility that the matrix models recently proposed to explain (almost) all the physics of M-theory may include the superstring theories that we know perturbatively. The ``1st quantized'' physical system of one IIA string seems…

High Energy Physics - Theory · Physics 2008-02-03 Lubos Motl
‹ Prev 1 2 3 10 Next ›