English
Related papers

Related papers: Segmenting DNA sequence into `words'

200 papers

Given a set of sequences, the distance between pairs of them helps us to find their similarity and derive structural relationship amongst them. For genomic sequences such measures make it possible to construct the evolution tree of…

Information Theory · Computer Science 2012-08-29 Sandeep Hosangadi

We consider a novel approach of measuring the homology of DNA sequences based of the variety of optimal alignments in the longest common subsequence sense. The proposed approach is compared with BLAST in measuring the homology of four…

Applications · Statistics 2012-10-16 Erik Hirmo , Jüri Lember , Heinrich Matzinger

We study electronic transport in long DNA chains using the tight-binding approach for a ladder-like model of DNA. We find insulating behavior with localizaton lengths xi ~ 25 in units of average base-pair seperation. Furthermore, we observe…

Soft Condensed Matter · Physics 2007-05-23 Antonio Rodriguez , Rudolf A. Roemer , Matthew S. Turner

In this paper, we present an optical computing method for string data alignment applicable to genome information analysis. By applying moire technique to spatial encoding patterns of deoxyribonucleic acid (DNA) sequences, association…

Most end-to-end speech recognition systems model text directly as a sequence of characters or sub-words. Current approaches to sub-word extraction only consider character sequence frequencies, which at times produce inferior sub-word…

Computation and Language · Computer Science 2019-02-22 Hainan Xu , Shuoyang Ding , Shinji Watanabe

We propose local segmentation of multiple sequences sharing a common time- or location-index, building upon the single sequence local segmentation methods of Niu and Zhang (2012) and Fang, Li and Siegmund (2016). We also propose reverse…

Statistics Theory · Mathematics 2022-01-10 Hock-Peng Chan , Hao Chen

Code generation maps a program description to executable source code in a programming language. Existing approaches mainly rely on a recurrent neural network (RNN) as the decoder. However, we find that a program contains significantly more…

Machine Learning · Computer Science 2018-11-19 Zeyu Sun , Qihao Zhu , Lili Mou , Yingfei Xiong , Ge Li , Lu Zhang

The rate of occurrence of words is not uniform but varies from document to document. Despite this observation, parameters for conventional n-gram language models are usually derived using the assumption of a constant word rate. In this…

Computation and Language · Computer Science 2007-05-23 Yoshihiko Gotoh , Steve Renals

Segmentation remains an important preprocessing step both in languages where "words" or other important syntactic/semantic units (like morphemes) are not clearly delineated by white space, as well as when dealing with continuous speech…

Computation and Language · Computer Science 2021-09-07 C. M. Downey , Fei Xia , Gina-Anne Levow , Shane Steinert-Threlkeld

This paper introduces Dynamic Programming Encoding (DPE), a new segmentation algorithm for tokenizing sentences into subword units. We view the subword segmentation of output sentences as a latent variable that should be marginalized out…

Computation and Language · Computer Science 2020-08-04 Xuanli He , Gholamreza Haffari , Mohammad Norouzi

Next-generation sequencing technologies generate millions of short sequence reads, which are usually aligned to a reference genome. In many applications, the key information required for downstream analysis is the number of reads mapping to…

Genomics · Quantitative Biology 2016-07-26 Yang Liao , Gordon K Smyth , Wei Shi

Tokenization is a crucial step in processing protein sequences for machine learning models, as proteins are complex sequences of amino acids that require meaningful segmentation to capture their functional and structural properties.…

Computation and Language · Computer Science 2024-11-27 Burak Suyunu , Enes Taylan , Arzucan Özgür

Linguistic analysis of protein sequences is an underexploited technique. Here, we capitalize on the concept of the lipogram to characterize sequences at the proteome levels. A lipogram is a literary composition which omits one or more…

Quantitative Methods · Quantitative Biology 2017-07-31 Jason Laurie , Amit K Chattopadhyay , Darren R Flower

This study proposes a data condensation method for multivariate kernel density estimation by genetic algorithm. First, our proposed algorithm generates multiple subsamples of a given size with replacement from the original sample. The…

Methodology · Statistics 2022-03-04 Kiheiji Nishida

Statistics about n-grams (i.e., sequences of contiguous words or other tokens in text documents or other string data) are an important building block in information retrieval and natural language processing. In this work, we study how…

Information Retrieval · Computer Science 2012-07-19 Klaus Berberich , Srikanta Bedathur

Background: With the fast development of next generation sequencing technologies, increasing numbers of genomes are being de novo sequenced and assembled. However, most are in fragmental and incomplete draft status, and thus it is often…

Genomics · Quantitative Biology 2020-02-28 Binghang Liu , Yujian Shi , Jianying Yuan , Xuesong Hu , Hao Zhang , Nan Li , Zhenyu Li , Yanxiang Chen , Desheng Mu , Wei Fan

We derive theoretical upper and lower bounds on the maximum size of DNA codes of length n with constant GC-content w and minimum Hamming distance d, both with and without the additional constraint that the minimum Hamming distance between…

Combinatorics · Mathematics 2007-05-23 Oliver D. King

The complementary strands of DNA molecules can be separated when stretched apart by a force; the unzipping signal is correlated to the base content of the sequence but is affected by thermal and instrumental noise. We consider here the…

Biomolecules · Quantitative Biology 2015-05-13 Valentina Baldazzi , Serena Bradde , Simona Cocco , Enzo Marinari , Remi Monasson

Synthetic DNA approaches 227.5 exabytes per gram of storage density with stability over millennial timescales. Realising this capacity requires error-correction codes that recover data from substantial synthesis and sequencing errors.…

Information Theory · Computer Science 2026-04-23 James L. Banal

Composite DNA is a recent method to increase the base alphabet size in DNA-based data storage.This paper models synthesizing and sequencing of composite DNA and introduces coding techniques to correct substitutions, losses of entire…

Information Theory · Computer Science 2025-10-29 Frederik Walter , Omer Sabary , Antonia Wachter-Zeh , Eitan Yaakobi