English

Error-Correcting Codes for Labeled DNA Sequences

Information Theory 2025-11-04 v1 math.IT

Abstract

Labeling of DNA molecules is a fundamental technique for DNA visualization and analysis. This process was mathematically modeled in [1], where the received sequence indicates the positions of the used labels. In this work, we develop error correcting codes for labeled DNA sequences, establishing bounds and constructing explicit systematic encoders for single substitution, insertion, and deletion errors. We focus on two cases: (1) using the complete set of length-two labels and (2) using the minimal set of length-two labels that ensures the recovery of DNA sequences from their labeling for 'almost' all DNA sequences.

Keywords

Cite

@article{arxiv.2511.01280,
  title  = {Error-Correcting Codes for Labeled DNA Sequences},
  author = {Dganit Hanania and Eitan Yaakobi},
  journal= {arXiv preprint arXiv:2511.01280},
  year   = {2025}
}