English
Related papers

Related papers: AI-generated text boundary detection with RoFT

200 papers

With the rapid development and widespread application of Large Language Models (LLMs), the use of Machine-Generated Text (MGT) has become increasingly common, bringing with it potential risks, especially in terms of quality and integrity in…

Computation and Language · Computer Science 2024-04-02 Qihui Zhang , Chujie Gao , Dongping Chen , Yue Huang , Yixin Huang , Zhenyang Sun , Shilin Zhang , Weiye Li , Zhengyan Fu , Yao Wan , Lichao Sun

As LLMs increase in accessibility, LLM-generated texts have proliferated across several fields, such as scientific, academic, and creative writing. However, LLMs are not created equally; they may have different architectures and training…

Computation and Language · Computer Science 2024-12-11 Shantanu Thorat , Tianbao Yang

With the increasing integration of large language models (LLMs) into open-domain writing, detecting machine-generated text has become a critical task for ensuring content authenticity and trust. Existing approaches rely on statistical…

Computation and Language · Computer Science 2025-10-15 Siyuan Li , Aodu Wulianghai , Xi Lin , Guangyan Li , Xiang Chen , Jun Wu , Jianhua Li

Detecting AI-generated text is a difficult problem to begin with; detecting AI-generated text on social media is made even more difficult due to the short text length and informal, idiosyncratic language of the internet. It is nonetheless…

Computation and Language · Computer Science 2025-06-17 Hillary Dawkins , Kathleen C. Fraser , Svetlana Kiritchenko

In this paper we analyze features to classify human- and AI-generated text for English, French, German and Spanish and compare them across languages. We investigate two scenarios: (1) The detection of text generated by AI from scratch, and…

Computation and Language · Computer Science 2024-01-31 Kristina Schaaff , Tim Schlippe , Lorenz Mindner

As advanced modern systems like deep neural networks (DNNs) and generative AI continue to enhance their capabilities in producing convincing and realistic content, the need to distinguish between user-generated and machine generated content…

Computation and Language · Computer Science 2024-04-01 Yaqi Xie , Anjali Rawal , Yujing Cen , Dixuan Zhao , Sunil K Narang , Shanu Sushmita

Nowadays, powerful large language models (LLMs) such as ChatGPT have demonstrated revolutionary power in a variety of tasks. Consequently, the detection of machine-generated texts (MGTs) is becoming increasingly crucial as LLMs become more…

Cryptography and Security · Computer Science 2024-01-17 Xinlei He , Xinyue Shen , Zeyuan Chen , Michael Backes , Yang Zhang

Content creation has dramatically progressed with the rapid advancement of large language models like ChatGPT and Claude. While this progress has greatly enhanced various aspects of life and work, it has also negatively affected certain…

Computation and Language · Computer Science 2025-06-05 Yuchen Guo , Zhicheng Dou , Huy H. Nguyen , Ching-Chun Chang , Saku Sugawara , Isao Echizen

With the advent of fluent generative language models that can produce convincing utterances very similar to those written by humans, distinguishing whether a piece of text is machine-generated or human-written becomes more challenging and…

Computation and Language · Computer Science 2024-02-27 Niloofar Mireshghallah , Justus Mattern , Sicun Gao , Reza Shokri , Taylor Berg-Kirkpatrick

Large language models (LLMs) can produce text that closely resembles human writing. This capability raises concerns about misuse, including disinformation and content manipulation. Detecting AI-generated text is essential to maintain…

Computation and Language · Computer Science 2025-12-29 Md. Rakibul Islam , Most. Sharmin Sultana Samu , Md. Zahid Hossain , Farhad Uz Zaman , Md. Kamrozzaman Bhuiyan

In the age of advanced large language models (LLMs), the boundaries between human and AI-generated text are becoming increasingly blurred. We address the challenge of segmenting mixed-authorship text, that is identifying transition points…

Computation and Language · Computer Science 2026-01-06 L. D. M. S. Sai Teja , N. Siva Gopala Krishna , Ufaq Khan , Muhammad Haris Khan , Atul Mishra

The rise in malicious usage of large language models, such as fake content creation and academic plagiarism, has motivated the development of approaches that identify AI-generated text, including those based on watermarking or outlier…

Computation and Language · Computer Science 2023-10-19 Kalpesh Krishna , Yixiao Song , Marzena Karpinska , John Wieting , Mohit Iyyer

The advancement of large language models (LLMs) has made it difficult to differentiate human-written text from AI-generated text. Several AI-text detectors have been developed in response, which typically utilize a fixed global threshold…

Computation and Language · Computer Science 2026-02-03 Minseok Jung , Cynthia Fuertes Panizo , Liam Dugan , Yi R. , Fung , Pin-Yu Chen , Paul Pu Liang

In recent years, text generation tools utilizing Artificial Intelligence (AI) have occasionally been misused across various domains, such as generating student reports or creative writings. This issue prompts plagiarism detection services…

Computation and Language · Computer Science 2025-04-14 Ahmed K. Kadhim , Lei Jiao , Rishad Shafik , Ole-Christoffer Granmo

The volume of machine-generated content online has grown dramatically due to the widespread use of Large Language Models (LLMs), leading to new challenges for content moderation systems. Conventional content moderation classifiers, which…

Computation and Language · Computer Science 2026-05-26 Shaz Furniturewala , Arkaitz Zubiaga

Large language models (LLMs) are solidifying their position in the modern world as effective tools for the automatic generation of text. Their use is quickly becoming commonplace in fields such as education, healthcare, and scientific…

Computation and Language · Computer Science 2025-10-08 Luka Terčon , Kaja Dobrovoljc

Research on AI-generated text detection has presented a number of approaches to discern human from AI prose, some of which achieving high in-distribution performance. However, real-world applicability has stalled because their outputs are…

Artificial Intelligence · Computer Science 2026-05-28 Aldan Creo , Suraj Ranganath

Large language models are increasingly used for many applications. To prevent illicit use, it is desirable to be able to detect AI-generated text. Training and evaluation of such detectors critically depend on suitable benchmark datasets.…

Machine Learning · Computer Science 2025-11-13 Philipp Dingfelder , Christian Riess

Training-free AI text detection methods primarily rely on model log-probabilities, achieving strong performance through approaches like Binoculars and DNA-DetectLLM. However, these methods face a fundamental ceiling as models are optimized…

Computation and Language · Computer Science 2026-05-05 Priyadarshan Narayanasamy , Swastik Agrawal , Klint Faber , Fardina Fathmiul Alam

With the development of generative models like GPT-3, it is increasingly more challenging to differentiate generated texts from human-written ones. There is a large number of studies that have demonstrated good results in bot…

Computation and Language · Computer Science 2023-11-21 Vasilii Gromov , Quynh Nhu Dang