中文
相关论文

相关论文: GPT4o-Receipt: A Dataset and Human Study for AI-Ge…

200 篇论文

Prior studies have shown that distinguishing text generated by Large Language Models (LLMs) from human-written one is highly challenging for humans, and often no better than random guessing. To verify the generalizability of this finding…

Peer review is a critical process for ensuring the integrity of published scientific research. Confidence in this process is predicated on the assumption that experts in the relevant domain give careful consideration to the merits of…

计算与语言 · 计算机科学 2024-12-09 Sungduk Yu , Man Luo , Avinash Madasu , Vasudev Lal , Phillip Howard

In this paper, we study how well humans can detect text generated by commercial LLMs (GPT-4o, Claude, o1). We hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide…

计算与语言 · 计算机科学 2025-05-21 Jenna Russell , Marzena Karpinska , Mohit Iyyer

Large Language Models (LLMs) are capable of reproducing human-like inferences, including inferences about emotions and mental states, from text. Whether this capability extends beyond text to other modalities remains unclear. Humans possess…

We examine whether a leading AI system GPT4 understands text as well as humans do, first using a well-established standardized test of discourse comprehension. On this test, GPT4 performs slightly, but not statistically significantly,…

计算与语言 · 计算机科学 2025-01-22 Thomas R. Shultz , Jamie M. Wise , Ardavan Salehi Nobandegani

AI-generated content is becoming increasingly prevalent in the real world, leading to serious ethical and societal concerns. For instance, adversaries might exploit large multimodal models (LMMs) to create images that violate ethical or…

计算与语言 · 计算机科学 2025-04-14 Hongchao Fang , Yixin Liu , Jiangshu Du , Can Qin , Ran Xu , Feng Liu , Lichao Sun , Dongwon Lee , Lifu Huang , Wenpeng Yin

Context: Code reviews are crucial for software quality. Recent AI advances have allowed large language models (LLMs) to review and fix code; now, there are tools that perform these reviews. However, their reliability and accuracy have not…

软件工程 · 计算机科学 2025-05-27 Umut Cihan , Arda İçöz , Vahid Haratian , Eray Tüzün

Legal invoice review is a costly, inconsistent, and time-consuming process, traditionally performed by Legal Operations, Lawyers or Billing Specialists who scrutinise billing compliance line by line. This study presents the first empirical…

计算与语言 · 计算机科学 2025-04-07 Nick Whitehouse , Nicole Lincoln , Stephanie Yiu , Lizzie Catterson , Rivindu Perera

This study evaluates the performance of ChatGPT variants, GPT-3.5 and GPT-4, both with and without prompt engineering, against solely student work and a mixed category containing both student and GPT-4 contributions in university-level…

计算与语言 · 计算机科学 2024-10-08 Will Yeadon , Alex Peach , Craig P. Testrow

Much is promised in relation to AI-supported software development. However, there has been limited evaluation effort in the research domain aimed at validating the true utility of such techniques, especially when compared to human coding…

软件工程 · 计算机科学 2025-01-29 Sherlock A. Licorish , Ansh Bajpai , Chetan Arora , Fanyu Wang , Kla Tantithamthavorn

This study investigates whether individuals can learn to accurately discriminate between human-written and AI-produced texts when provided with immediate feedback, and if they can use this feedback to recalibrate their self-perceived…

计算与语言 · 计算机科学 2025-10-17 Jiří Milička , Anna Marklová , Ondřej Drobil , Eva Pospíšilová

Retrieval-augmented generation (RAG) enables large language models (LLMs) to generate answers with citations from source documents containing "ground truth", thereby reducing system hallucinations. A crucial factor in RAG evaluation is…

计算与语言 · 计算机科学 2025-04-22 Nandan Thakur , Ronak Pradeep , Shivani Upadhyay , Daniel Campos , Nick Craswell , Jimmy Lin

Large language models (LLMs) can extract information from veterinary electronic health records (EHRs), but performance differences between models, the effect of temperature settings, and the influence of text ambiguity have not been…

As AI systems increasingly evaluate other AI outputs, understanding their assessment behavior becomes crucial for preventing cascading biases. This study analyzes vision-language descriptions generated by NVIDIA's Describe Anything Model…

人工智能 · 计算机科学 2025-09-22 Sajjad Abdoli , Rudi Cilibrasi , Rima Al-Shikh

Generative AI (GenAI) holds significant promise for automating everyday image editing tasks, especially following the recent release of GPT-4o on March 25, 2025. However, what subjects do people most often want edited? What kinds of editing…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Mohammad Reza Taesiri , Brandon Collins , Logan Bolton , Viet Dac Lai , Franck Dernoncourt , Trung Bui , Anh Totti Nguyen

The advancing fluency of LLMs raises important questions about their ability to emulate complex human traits, including emotional expression and personality, across diverse linguistic and cultural contexts. This study investigates whether…

计算与语言 · 计算机科学 2026-03-25 Nasser A Alsadhan

We test the abilities of specialised deep neural networks like PersonalityMap as well as general LLMs like GPT-4o and Claude 3 Opus in understanding human personality. Specifically, we compare their ability to predict correlations between…

计算机与社会 · 计算机科学 2024-06-13 Philipp Schoenegger , Spencer Greenberg , Alexander Grishin , Joshua Lewis , Lucius Caviola

Recently, generative AIs like ChatGPT have become available to the wide public. These tools can for instance be used by students to generate essays or whole theses. But how does a teacher know whether a text is written by a student or an…

计算与语言 · 计算机科学 2023-11-14 Lorenz Mindner , Tim Schlippe , Kristina Schaaff

Generative AI now produces photorealistic portraits that circulate widely in social and newslike contexts. Human ability to distinguish real from synthetic faces is time-sensitive because image generators continue to improve while public…

人机交互 · 计算机科学 2026-03-26 Sunwhi Kim , Sunyul Kim

The rapid advancement of Large Language Models (LLMs) presents a significant challenge to academic integrity within computing education. As educators seek reliable detection methods, this paper evaluates the capacity of three prominent LLMs…

计算机与社会 · 计算机科学 2025-12-30 Christopher Burger , Karmece Talley , Christina Trotter
‹ 上一页 1 2 3 10 下一页 ›