English
Related papers

Related papers: Detecting Machine-Generated Long-Form Content with…

200 papers

The proliferation of large language models (LLMs) has significantly transformed the digital information landscape, making it increasingly challenging to distinguish between human-written and LLM-generated content. Detecting LLM-generated…

Computation and Language · Computer Science 2025-06-30 Minjia Mao , Dongjun Wei , Xiao Fang , Michael Chau

This paper presents a novel approach for detecting ChatGPT-generated vs. human-written text using language models. To this end, we first collected and released a pre-processed dataset named OpenGPTText, which consists of rephrased content…

Computation and Language · Computer Science 2023-05-19 Yutian Chen , Hao Kang , Vivian Zhai , Liangze Li , Rita Singh , Bhiksha Raj

Generative models, especially large language models (LLMs), have shown remarkable progress in producing text that appears human-like. However, they often exhibit patterns that make their output easier to detect than text written by humans.…

Computation and Language · Computer Science 2026-01-06 Hadi Mohammadi , Anastasia Giachanou , Daniel L. Oberski , Ayoub Bagheri

ChatGPT has become a global sensation. As ChatGPT and other Large Language Models (LLMs) emerge, concerns of misusing them in various ways increase, such as disseminating fake news, plagiarism, manipulating public opinion, cheating, and…

Machine Learning · Computer Science 2023-04-06 Alessandro Pegoraro , Kavita Kumari , Hossein Fereidooni , Ahmad-Reza Sadeghi

While large language models (LLMs) have made significant strides in generating coherent and contextually relevant text, they often function as opaque black boxes, trained on vast unlabeled datasets with statistical objectives, lacking an…

Computation and Language · Computer Science 2025-03-03 Yingbing Huang , Deming Chen , Abhishek K. Umrawal

Large Language Models (LLMs) have demonstrated remarkable capabilities in generating text that closely resembles human writing across a wide range of styles and genres. However, such capabilities are prone to potential misuse, such as fake…

Computation and Language · Computer Science 2025-05-20 Harika Abburi , Sanmitra Bhattacharya , Edward Bowen , Nirmala Pudota

Recent advances in natural language processing (NLP) may enable artificial intelligence (AI) models to generate writing that is identical to human written form in the future. This might have profound ethical, legal, and social…

Machine Learning · Computer Science 2024-04-17 Nuzhat Prova

Machine-generated texts (MGTs) pose risks such as disinformation and phishing, underscoring the need for reliable detection. Metric-based methods, which extract statistically distinguishable features of MGTs, are often more practical than…

Computation and Language · Computer Science 2026-05-18 Chenwang Wu , Yiuming Cheung , Bo Han , Shuhai Zhang , Defu Lian

With the increasing quality and spread of LLM assistants, the amount of generated content is growing rapidly. In many cases and tasks, such texts are already indistinguishable from those written by humans, and the quality of generation…

Computation and Language · Computer Science 2026-04-15 Irina Tolstykh , Aleksandra Tsybina , Sergey Yakubson , Aleksandr Gordeev , Vladimir Dokholyan , Maksim Kuprashevich

Detecting AI-generated text is a difficult problem to begin with; detecting AI-generated text on social media is made even more difficult due to the short text length and informal, idiosyncratic language of the internet. It is nonetheless…

Computation and Language · Computer Science 2025-06-17 Hillary Dawkins , Kathleen C. Fraser , Svetlana Kiritchenko

The potential misuse of ChatGPT and other Large Language Models (LLMs) has raised concerns regarding the dissemination of false information, plagiarism, academic dishonesty, and fraudulent activities. Consequently, distinguishing between…

Cryptography and Security · Computer Science 2023-11-10 Kavita Kumari , Alessandro Pegoraro , Hossein Fereidooni , Ahmad-Reza Sadeghi

Detecting Large Language Model (LLM)-generated code is a growing challenge with implications for security, intellectual property, and academic integrity. We investigate the role of conditional probability distributions in improving…

Computation and Language · Computer Science 2025-06-09 Maor Ashkenazi , Ofir Brenner , Tal Furman Shohet , Eran Treister

With the rise of advanced natural language models like GPT, distinguishing between human-written and GPT-generated text has become increasingly challenging and crucial across various domains, including academia. The long-standing issue of…

Computation and Language · Computer Science 2025-10-14 A. Selvioğlu , V. Adanova , M. Atagoziev

With the launch of ChatGPT, large language models (LLMs) have attracted global attention. In the realm of article writing, LLMs have witnessed extensive utilization, giving rise to concerns related to intellectual property protection,…

Computation and Language · Computer Science 2024-06-14 Ying Zhou , Ben He , Le Sun

Large language models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their ability to generate human-like text has raised concerns about potential misuse. This underscores the need for reliable and effective…

Computation and Language · Computer Science 2026-04-24 Runheng Liu , Heyan Huang , Xingchen Xiao , Zhijing Wu

Sensitive information detection is crucial in content moderation to maintain safe online communities. Assisting in this traditionally manual process could relieve human moderators from overwhelming and tedious tasks, allowing them to focus…

Following the universal availability of generative AI systems with the release of ChatGPT, automatic detection of deceptive text created by Large Language Models has focused on domains such as academic plagiarism and "fake news". However,…

Computation and Language · Computer Science 2024-12-23 Andrea Cristina McGlinchey , Peter J Barclay

The rapid proliferation of large language models (LLMs) has increased the volume of machine-generated texts (MGTs) and blurred text authorship in various domains. However, most existing MGT benchmarks include single-author texts…

Computation and Language · Computer Science 2025-03-18 Ekaterina Artemova , Jason Lucas , Saranya Venkatraman , Jooyoung Lee , Sergei Tilga , Adaku Uchendu , Vladislav Mikhailov

Neural text detectors are models trained to detect whether a given text was generated by a language model or written by a human. In this paper, we investigate three simple and resource-efficient strategies (parameter tweaking, prompt…

Computation and Language · Computer Science 2023-11-06 Vitalii Fishchuk , Daniel Braun

Large language models (LLMs) are increasingly used in daily applications, from content generation to code writing, where each interaction treats the model as stateless, generating responses independently without memory. Yet human writing is…

Computation and Language · Computer Science 2026-04-15 Zhanwei Cao , YeoJin Go , Yifan Hu , Shanu Sushmita