English

SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection

Computation and Language 2025-07-02 v2

Abstract

The increasing capability of large language models (LLMs) to generate synthetic content has heightened concerns about their misuse, driving the development of Machine-Generated Text (MGT) detection models. However, these detectors face significant challenges due to the lack of high-quality synthetic datasets for training. To address this issue, we propose SPADE, a structured framework for detecting synthetic dialogues using prompt-based positive and negative samples. Our proposed methods yield 14 new dialogue datasets, which we benchmark against eight MGT detection models. The results demonstrate improved generalization performance when utilizing a mixed dataset produced by proposed augmentation frameworks, offering a practical approach to enhancing LLM application security. Considering that real-world agents lack knowledge of future opponent utterances, we simulate online dialogue detection and examine the relationship between chat history length and detection accuracy. Our open-source datasets, code and prompts can be downloaded from https://github.com/AngieYYF/SPADE-customer-service-dialogue.

Keywords

Cite

@article{arxiv.2503.15044,
  title  = {SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection},
  author = {Haoyi Li and Angela Yifei Yuan and Soyeon Caren Han and Christopher Leckie},
  journal= {arXiv preprint arXiv:2503.15044},
  year   = {2025}
}

Comments

ACL LLMSEC

R2 v1 2026-06-28T22:26:34.460Z