English
Related papers

Related papers: GPT4o-Receipt: A Dataset and Human Study for AI-Ge…

200 papers

The development of Generative AI Large Language Models (LLMs) raised the alarm regarding identifying content produced through generative AI or humans. In one case, issues arise when students heavily rely on such tools in a manner that can…

Computation and Language · Computer Science 2025-01-07 Ayat Najjar , Huthaifa I. Ashqar , Omar Darwish , Eman Hammad

This study examines the feasibility and potential advantages of using large language models, in particular GPT-4o, to perform partial credit grading of large numbers of student written responses to introductory level physics problems.…

Physics Education · Physics 2025-08-21 Zhongzhou Chen , Tong Wan

Face image synthesis has progressed beyond the point at which humans can effectively distinguish authentic faces from synthetically generated ones. Recently developed synthetic face image detectors boast "better-than-human" discriminative…

Computer Vision and Pattern Recognition · Computer Science 2022-08-24 Aidan Boyd , Patrick Tinsley , Kevin Bowyer , Adam Czajka

As AI writing tools become widespread, we need to understand how both humans and machines evaluate literary style, a domain where objective standards are elusive and judgments are inherently subjective. We conducted controlled experiments…

Artificial Intelligence · Computer Science 2025-10-13 Wouter Haverals , Meredith Martin

How do Large Language Models understand moral dimensions compared to humans? This first large-scale Bayesian evaluation of market-leading language models provides the answer. In contrast to prior work using deterministic ground truth…

Computation and Language · Computer Science 2025-11-24 Maciej Skorski , Alina Landowska

Deep learning is closing the gap with human vision on several object recognition benchmarks. Here we investigate this gap for challenging images where objects are seen in unusual poses. We find that humans excel at recognizing objects in…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Netta Ollikka , Amro Abbas , Andrea Perin , Markku Kilpeläinen , Stéphane Deny

We investigate the multilingual and multimodal performance of a large language model-based artificial intelligence (AI) system, GPT-4o, using a diverse set of physics concept inventories spanning multiple languages and subject categories.…

Physics Education · Physics 2025-07-14 Gerd Kortemeyer , Marina Babayeva , Giulia Polverini , Ralf Widenhorn , Bor Gregorcic

Generative artificial intelligence tools, like ChatGPT, are an increasingly utilized resource among computational social scientists. Nevertheless, there remains space for improved understanding of the performance of ChatGPT in complex tasks…

Computation and Language · Computer Science 2025-12-02 Breanna E. Green , Ashley L. Shea , Pengfei Zhao , Drew B. Margolin

Heuristic evaluation is a widely used method in Human-Computer Interaction (HCI) to inspect interfaces and identify issues based on heuristics. Recently, Large Language Models (LLMs), such as GPT-4o, have been applied in HCI to assist in…

Human-Computer Interaction · Computer Science 2026-05-12 Guilherme Guerino , Luiz Rodrigues , Bruna Capeleti , Rafael Ferreira Mello , André Freire , Luciana Zaina

This study investigates whether large language models, specifically GPT4, can match human capabilities in analogical reasoning within strategic decision making contexts. Using a novel experimental design involving source to target matching,…

Artificial Intelligence · Computer Science 2025-05-02 Phanish Puranam , Prothit Sen , Maciej Workiewicz

As dialogue systems and chatbots increasingly integrate into everyday interactions, the need for efficient and accurate evaluation methods becomes paramount. This study explores the comparative performance of human and AI assessments across…

Computation and Language · Computer Science 2024-09-11 Ike Ebubechukwu , Johane Takeuchi , Antonello Ceravola , Frank Joublin

The Uniform Information Density (UID) principle posits that humans prefer to spread information evenly during language production. We examine if this UID principle can help capture differences between Large Language Models (LLMs)-generated…

Computation and Language · Computer Science 2024-04-05 Saranya Venkatraman , Adaku Uchendu , Dongwon Lee

Large language models are increasingly used to curate bibliographies, raising the question: are their reference lists distinguishable from human ones? We build paired citation graphs, ground truth and GPT-4o-generated (from parametric…

Machine Learning · Computer Science 2026-01-29 Melika Mobini , Vincent Holst , Floriano Tori , Andres Algaba , Vincent Ginis

This paper investigates why recent generative AI models outperform humans in data visualization knowledge tasks. Through systematic comparative analysis of responses to visualization questions, we find that differences exist between two…

Human-Computer Interaction · Computer Science 2025-08-05 Yongsu Ahn , Nam Wook Kim

The growing capability of large language models to produce fluent, contextually coherent text has created mounting pressure on the systems and institutions responsible for ensuring the authenticity of digital content. Advanced generative…

This study evaluates the performance of large language models (LLMs) and the HINT model in predicting clinical trial outcomes, focusing on metrics including Balanced Accuracy, Matthews Correlation Coefficient (MCC), Recall, and Specificity.…

Machine Learning · Computer Science 2025-03-19 Shuyi Jin , Lu Chen , Hongru Ding , Meijie Wang , Lun Yu

As LLMs become increasingly proficient at producing human-like responses, there has been a rise of academic and industrial pursuits dedicated to flagging a given piece of text as "human" or "AI". Most of these pursuits involve modern NLP…

Artificial Intelligence · Computer Science 2024-09-10 Prathamesh Dinesh Joshi , Sahil Pocker , Raj Abhijit Dandekar , Rajat Dandekar , Sreedath Panat

The misuse of generative AI in online disinformation campaigns highlights the urgent need for transparent and explainable detection systems. In this work, we investigate how detectors for AI-generated images can be more effective in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Silvia Poletti , Justin Ilyes , Marcel Hasenbalg , David Fischinger , Martin Boyer

The rapid proliferation of large language models (LLMs) has created an urgent need for robust and generalizable detectors of machine-generated text. Existing benchmarks typically evaluate a single detector on a single dataset under ideal…

Computation and Language · Computer Science 2026-03-19 Madhav S. Baidya , S. S. Baidya , Chirag Chawla

As large language models (LLMs) generate text that increasingly resembles human writing, the subtle cues that distinguish AI-generated content from human-written content become increasingly challenging to capture. Reliance on…

Computation and Language · Computer Science 2026-04-16 Xiao Pu , Zepeng Cheng , Lin Yuan , Yu Wu , Xiuli Bi