English
Related papers

Related papers: Revisiting Waiting Times in DNA evolution

200 papers

Multiple hypothesis testing is a significant problem in nearly all neuroimaging studies. In order to correct for this phenomena, we require a reliable estimate of the Family-Wise Error Rate (FWER). The well known Bonferroni correction…

Computation · Statistics 2015-02-17 Chris Hinrichs , Vamsi K Ithapu , Qinyuan Sun , Sterling C Johnson , Vikas Singh

In Bayesian phylogenetics, our goal is to estimate the posterior distribution over phylogenetic trees. Markov chain Monte Carlo methods are widely used to approximate the phylogenetic posterior distributions. For large-scale sequence data,…

Methodology · Statistics 2026-05-12 Wentao Yu , Shijia Wang

A correlation between karyotype diversity and species richness was first observed in mammals in 1980, and subsequently confirmed after controlling for phylogenetic signal. The correlation was attributed to submicroscopic factors, presumably…

Populations and Evolution · Quantitative Biology 2025-11-18 John Herrick

Next-generation sequencing techniques have facilitated a large scale analysis of human genetic variation. Despite the advances in sequencing speeds, the computational discovery of structural variants is not yet standard. It is likely that…

Our results demonstrated that a previously reported protein name co-occurrence method (5-mention PubGene) which was not based on a hypothesis testing framework, it is generally statistically more significant than the 99th percentile of…

Digital Libraries · Computer Science 2009-01-05 Maurice HT Ling , Christophe Lefevre , Kevin R. Nicholas

We consider a branching population where individuals live and reproduce independently. Their lifetimes are i.i.d. and they give birth at a constant rate b. The genealogical tree spanned by this process is called a splitting tree, and the…

Probability · Mathematics 2016-09-05 Nicolas Champagnat , Benoît Henry

Genome assembly, the process of reconstructing a long genetic sequence by aligning and merging short fragments, or reads, is known to be NP-hard, either as a version of the shortest common superstring problem or in a Hamiltonian-cycle…

Statistical Mechanics · Physics 2024-03-12 L. A. Fernandez , V. Martin-Mayor , D. Yllanes

The Luria-Delbr\"uck model is a classic model of population dynamics with random mutations, that has been used historically to prove that random mutations drive evolution. In typical scenarios, the relevant mutation rate is exceedingly…

Biological Physics · Physics 2024-02-22 Deng Pan , Jie Lin , Ariel Amir

The discrete distribution of the length of longest increasing subsequences in random permutations of $n$ integers is deeply related to random matrix theory. In a seminal work, Baik, Deift and Johansson provided an asymptotics in terms of…

Combinatorics · Mathematics 2024-06-21 Folkmar Bornemann

Monte Carlo methods can provide accurate p-value estimates of word counting test statistics and are easy to implement. They are especially attractive when an asymptotic theory is absent or when either the search sequence or the word pattern…

Applications · Statistics 2008-12-01 Hock Peng Chan , Nancy R. Zhang , Louis H. Y. Chen

{\it Transcription} is the process whereby RNA molecules are polymerized by molecular machines, called RNA polymerase (RNAP), using the corresponding DNA as the template. Recent {\it in-vivo} experiments with single cells have established…

Biological Physics · Physics 2009-11-13 Tripti Tripathi , Debashish Chowdhury

DNA sequencing is the process of determining the exact order of the nucleotide bases of an individual's genome in order to catalogue sequence variation and understand its biological implications. Whole-genome sequencing techniques produce…

Data Structures and Algorithms · Computer Science 2015-09-18 Ljiljana Brankovic , Costas S. Iliopoulos , Ritu Kundu , Manal Mohamed , Solon P. Pissis , Fatima Vayani

Sampling is a common strategy for generating text from probabilistic models, yet standard ancestral sampling often results in text that is incoherent or ungrammatical. To alleviate this issue, various modifications to a model's sampling…

Computation and Language · Computer Science 2024-01-08 Clara Meister , Tiago Pimentel , Luca Malagutti , Ethan G. Wilcox , Ryan Cotterell

In many biochemical processes, proteins bound to DNA at distant sites are brought into close proximity by loops in the underlying DNA. For example, the function of some gene-regulatory proteins depends on such DNA looping interactions. We…

Quantitative Methods · Quantitative Biology 2009-11-13 John F Beausang , Philip C Nelson

Although real-world text datasets, such as DNA sequences, are far from being uniformly random, average-case string searching algorithms perform significantly better than worst-case ones in most applications of interest. In this paper, we…

Data Structures and Algorithms · Computer Science 2018-01-16 Lorraine A. K. Ayad , Panagiotis Charalampopoulos , Costas S. Iliopoulos , Solon P. Pissis

This paper establishes formal mathematical foundations linking Chaos Game Representations (CGR) of DNA sequences to their underlying $k$-mer frequencies. We prove that the Frequency CGR (FCGR) of order $k$ is mathematically equivalent to a…

Formal Languages and Automata Theory · Computer Science 2025-07-01 Haoze He , Lila Kari , Pablo Millan Arias

Transformers are neural networks that revolutionized natural language processing and machine learning. They process sequences of inputs, like words, using a mechanism called self-attention, which is trained via masked language modeling…

Disordered Systems and Neural Networks · Physics 2024-04-17 Riccardo Rende , Federica Gerace , Alessandro Laio , Sebastian Goldt

Large language models (LLMs) trained on text demonstrated remarkable results on natural language processing (NLP) tasks. These models have been adapted to decipher the language of DNA, where sequences of nucleotides act as "words" that…

We present some results of simulations of population growth and evolution, using the standard asexual Penna model, with individuals characterized by a string of bits representing a genome containing some possible mutations. After about…

Populations and Evolution · Quantitative Biology 2009-11-11 Mikolaj Sitarz , Andrzej Z. Maksymowicz

This paper deals with the two fundamental problems concerning the handling of large n-gram language models: indexing, that is compressing the n-gram strings and associated satellite data without compromising their retrieval speed; and…

Information Retrieval · Computer Science 2022-02-08 Giulio Ermanno Pibiri , Rossano Venturini
‹ Prev 1 8 9 10 Next ›