English

Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness

Sound 2026-07-23 v1

Abstract

Neural audio codecs are challenging transformations for audio watermarking because they re-encode, quantize, and resynthesize speech. This paper investigates continuous latent-space watermarking for codec robustness. Instead of adding a watermark only to the waveform or spectrogram, we embed a 32-bit message into the continuous latent representation of a codec-like speech autoencoder. The pipeline uses a SEANet-style encoder-decoder, a Conformer-based message embedder, RVQ-guided latent decomposition, and a latent-domain detector trained under signal-processing and neural-codec transformations. Rather than proposing a final universal watermarking baseline, we characterize the trade-offs that appear when the watermark carrier is moved before neural decoding. On 48 kHz speech, EnCodec-aware training improves EnCodec-24k bit accuracy from 78.8% to 95.6% and 97.1%, while PESQ decreases from 3.727 to 3.514 and 3.427.

Cite

@article{arxiv.2607.21132,
  title  = {Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness},
  author = {Zi Hu and Houmin Sun and Linxi Li and Yechen Wang and Liwei Jin and Carsten Maple and Ming Li},
  journal= {arXiv preprint arXiv:2607.21132},
  year   = {2026}
}

Comments

7 pages. Submitted to IEEE Spoken Language Technology Workshop (SLT)