English

Frequency Coding over Noisy Sampling

Information Theory 2026-08-01 v1

Abstract

DNA molecules are so small that it might be practical to use their frequency vectors to encode messages. More precisely, a sender can inject MXM_X copies of the string X=X = CATCATCAT into a pool and the receiver can recover MXM_X by sequencing the pool. There are, however, two sources of uncertainty: (a) MXM_X is usually too big to be counted exactly, but is estimated by sampling. (b) The DNA sequencer could be noisy; it may have difficulty distinguishing CATCATCAT from CATGATCAT. Recently, Tamir, Weinberger, and Guill\'en i F\`abregas clarified the amount of information the frequency vector can carry under (a). They showed that each string can carry about log4R\log_4 R bits, where RR is the average number of times each string is read. They also showed that log4R\log_4 R bits can be achieved by a low-complexity uncoded scheme under the condition that there are at least R\sqrt R distinct strings. In this paper, we show that a low-complexity coded scheme can achieve the same log4R\log_4 R bits unconditionally. We then generalize the scheme to handle sequencing noise, (b), and show that the noise penalizes the total number of bits by log2detW\log_2 \det W, together with a linear term due to the use of Fourier transforms in our proof. The former penalty log2detW\log_2 \det W is asymptotically the same as that obtained by Gerzon, Shomorony, and Weinberger; our scheme trades a small amount of rate for practical complexity.

Cite

@article{arxiv.2608.00539,
  title  = {Frequency Coding over Noisy Sampling},
  author = {Bo-Yu Su and Hsin-Po Wang and Venkatesan Guruswami},
  journal= {arXiv preprint arXiv:2608.00539},
  year   = {2026}
}

Comments

7 pages