English
Related papers

Related papers: Haplotype Assembly: An Information Theoretic View

200 papers

Motivation: DNA data is transcribed into single-stranded RNA, which folds into specific molecular structures. In this paper we pose the question to what extent sequence- and structure-information correlate. We view this correlation as…

Combinatorics · Mathematics 2016-08-23 Christopher Barrett , Fenix W. Huang , Christian M. Reidys

Decoding sequences that stem from multiple transmissions of a codeword over an insertion, deletion, and substitution channel is a critical component of efficient deoxyribonucleic acid (DNA) data storage systems. In this paper, we consider a…

Information Theory · Computer Science 2022-09-13 Issam Maarouf , Andreas Lenz , Lorenz Welter , Antonia Wachter-Zeh , Eirik Rosnes , Alexandre Graell i Amat

In this paper, we study the Random Access Problem in DNA storage, which addresses the challenge of retrieving a specific information strand from a DNA-based storage system. In this framework, the data is represented by $k$ information…

Information Theory · Computer Science 2025-08-26 Avital Boruchovsky , Ohad Elishco , Ryan Gabrys , Anina Gruica , Itzhak Tamo , Eitan Yaakobi

Short tandem repeats (STRs) and single nucleotide polymorphisms (SNPs) are two kinds of commonly used markers in Y chromosome studies of forensic and population genetics. There has been increasing interest in the cost saving strategy by…

Populations and Evolution · Quantitative Biology 2013-10-22 Chuan-Chao Wang , Ling-Xiang Wang , Rukesh Shrestha , Shaoqing Wen , Manfei Zhang , Xinzhu Tong , Li Jin , Hui Li

A new family of codes, called clustering-correcting codes, is presented in this paper. This family of codes is motivated by the special structure of data that is stored in DNA-based storage systems. The data stored in these systems has the…

Information Theory · Computer Science 2019-03-12 Tal Shinkar , Eitan Yaakobi , Andreas Lenz , Antonia Wachter-Zeh

High-throughput shotgun sequence data makes it possible in principle to accurately estimate population genetic parameters without confounding by SNP ascertainment bias. One such statistic of interest is the proportion of heterozygous sites…

Populations and Evolution · Quantitative Biology 2012-12-18 Katarzyna Bryc , Nick Patterson , David Reich

We study the relation between the persistent homology and the spectral sequence of a filtered chain complex over a field. Our method is based on a decomposition of the persistent homology. We demonstrate that, under fairly general…

Algebraic Topology · Mathematics 2024-03-25 Peiqi Yang , Yingfeng Hu , Hao Wu

We describe a strategy for constructing codes for DNA-based information storage by serial composition of weighted finite-state transducers. The resulting state machines can integrate correction of substitution errors; synchronization by…

Information Theory · Computer Science 2016-11-18 Ian Holmes

The problem of storing large amounts of information safely for a long period of time has become essential. One of the most promising new data storage mediums are the polymer-based data storage systems, like the DNA-storage system. These…

Information Theory · Computer Science 2025-04-21 Ville Junnila , Tero Laihonen , Tuomo Lehtilä

While many short read assemblers attempt to simplify the de Brujin graph by identifying and resolving variant-induced bubbles to produce a haploid mosaic result, this approach is only viable when variants are relatively rare and the bubbles…

Genomics · Quantitative Biology 2017-03-30 Eugene Goltsman , Isaac Ho , Daniel Rokhsar

Recent experiments have demonstrated the feasibility of storing digital information in macromolecules such as DNA and protein. However, the DNA storage channel is prone to errors such as deletions, insertions, and substitutions. During the…

Information Theory · Computer Science 2024-10-22 Aryan Abbasian , Mahtab Mirmohseni , Masoumeh Nasiri Kenari

DNA data storage systems encode digital data into DNA strands, enabling dense and durable storage. Efficient data retrieval depends on coverage depth, a key performance metric. We study the random access coverage depth problem and focus on…

Information Theory · Computer Science 2025-07-29 Şeyma Bodur , Stefano Lia , Hiram H. López , Rati Ludhani , Alberto Ravagnani , Lisa Seccia

Motivation: Transcriptome sequencing has long been the favored method for quickly and inexpensively obtaining the sequences for a large number of genes from an organism with no reference genome. With the rapidly increasing throughputs and…

Pyrosequencing is among the emerging sequencing techniques, capable of generating upto 100,000 overlapping reads in a single run. This technique is much faster and cheaper than the existing state of the art sequencing technique such as…

Genomics · Quantitative Biology 2016-09-08 Fahad Saeed , Ashfaq Khokhar , Osvaldo Zagordi , Niko Beerenwinkel

We formulate genome assembly problem as an optimization problem in which the objective function is the likelihood of the assembly given the reads.

Computational Engineering, Finance, and Science · Computer Science 2016-04-08 Mohammadreza Ghodsi

The human genome contains repetitive DNA at different level of sequence length, number and dispersion. Highly repetitive DNA is particularly rich in homo-- and di--nucleotide repeats, while middle repetitive DNA is rich of families of…

Genomics · Quantitative Biology 2009-11-11 Francesco Piazza , Pietro Lio

In this paper, we solve the sample shortage problem in the human parsing task. We begin with the self-learning strategy, which generates pseudo-labels for unlabeled data to retrain the model. However, directly using noisy pseudo-labels will…

Computer Vision and Pattern Recognition · Computer Science 2020-04-20 Tao Li , Zhiyuan Liang , Sanyuan Zhao , Jiahao Gong , Jianbing Shen

Complexity metrics and machine learning (ML) models have been utilized to analyze the lengths of segmental genomic entities like: exons, introns, intergenic and repeat/unique DNA sequences, in each of the 22 human chromosomes. The purpose…

The high-throughput short-reads RNA-seq protocols often produce paired-end reads, with the middle portion of the fragments being unsequenced. We explore if the full-length fragments can be computationally reconstructed from the sequenced…

Genomics · Quantitative Biology 2023-10-06 Xiang Li , Mingfu Shao

Owing to its immense storage density and durability, DNA has emerged as a promising storage medium. However, due to technological constraints, data can only be written onto many short DNA molecules called data blocks that are stored in an…

Information Theory · Computer Science 2024-03-26 Shubhransh Singhvi , Charchit Gupta , Avital Boruchovsky , Yuval Goldberg , Han Mao Kiah , Eitan Yaakobi