English

HC3 Plus: A Semantic-Invariant Human ChatGPT Comparison Corpus

Computation and Language 2024-10-10 v4 Artificial Intelligence

Abstract

ChatGPT has garnered significant interest due to its impressive performance; however, there is growing concern about its potential risks, particularly in the detection of AI-generated content (AIGC), which is often challenging for untrained individuals to identify. Current datasets used for detecting ChatGPT-generated text primarily focus on question-answering tasks, often overlooking tasks with semantic-invariant properties, such as summarization, translation, and paraphrasing. In this paper, we demonstrate that detecting model-generated text in semantic-invariant tasks is more challenging. To address this gap, we introduce a more extensive and comprehensive dataset that incorporates a wider range of tasks than previous work, including those with semantic-invariant properties. In addition, instruction fine-tuning has demonstrated superior performance across various tasks. In this paper, we explore the use of instruction fine-tuning models for detecting text generated by ChatGPT.

Keywords

Cite

@article{arxiv.2309.02731,
  title  = {HC3 Plus: A Semantic-Invariant Human ChatGPT Comparison Corpus},
  author = {Zhenpeng Su and Xing Wu and Wei Zhou and Guangyuan Ma and Songlin Hu},
  journal= {arXiv preprint arXiv:2309.02731},
  year   = {2024}
}

Comments

This paper has been accepted by CIKM2023 workshop

R2 v1 2026-06-28T12:13:52.890Z