English

Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework

Cryptography and Security 2026-04-14 v2 Computation and Language

Abstract

While watermarking serves as a critical mechanism for LLM provenance, existing secret-key schemes tightly couple detection with injection, requiring access to keys or provider-side scheme-specific detectors for verification. This dependency creates a fundamental barrier for real-world governance, as independent auditing becomes impossible without compromising model security or relying on the opaque claims of service providers. To resolve this dilemma, we introduce TTP-Detect, a pioneering black-box framework designed for non-intrusive, third-party watermark verification. By decoupling detection from injection, TTP-Detect reframes verification as a relative hypothesis testing problem. It employs a proxy model to amplify watermark-relevant signals and a suite of complementary relative measurements to assess the alignment of the query text with watermarked distributions. Extensive experiments across representative watermarking schemes, datasets and models demonstrate that TTP-Detect achieves superior detection performance and robustness against diverse attacks.

Keywords

Cite

@article{arxiv.2603.14968,
  title  = {Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework},
  author = {Zhuoshang Wang and Yubing Ren and Yanan Cao and Fang Fang and Xiaoxue Li and Li Guo},
  journal= {arXiv preprint arXiv:2603.14968},
  year   = {2026}
}

Comments

Accepted to ACL 2026 Findings

R2 v1 2026-07-01T11:21:49.320Z