中文
相关论文

相关论文: People cannot distinguish GPT-4 from a human in a …

200 篇论文

Large language models (LLMs) are increasingly used in the social sciences to simulate human behavior, based on the assumption that they can generate realistic, human-like text. Yet this assumption remains largely untested. Existing…

计算与语言 · 计算机科学 2025-11-26 Nicolò Pagan , Petter Törnberg , Christopher A. Bail , Anikó Hannák , Christopher Barrie

We aim to understand how people assess human likeness in navigation produced by people and artificially intelligent (AI) agents in a video game. To this end, we propose a novel AI agent with the goal of generating more human-like behavior.…

Turing test was originally proposed to examine whether machine's behavior is indistinguishable from a human. The most popular and practical Turing test is CAPTCHA, which is to discriminate algorithm from human by offering recognition-alike…

计算机视觉与模式识别 · 计算机科学 2019-04-23 Jiaming Zhang , Jitao Sang , Kaiyuan Xu , Shangxi Wu , Yongli Hu , Yanfeng Sun , Jian Yu

Large-scale AI models such as GPT-4 have accelerated the deployment of artificial intelligence across critical domains including law, healthcare, and finance, raising urgent questions about trust and transparency. This study investigates…

人工智能 · 计算机科学 2025-10-20 Allen Daniel Sunny

GPT-4 is often heralded as a leading commercial AI offering, sparking debates over its potential as a steppingstone toward artificial general intelligence. But does it possess consciousness? This paper investigates this key question using…

人工智能 · 计算机科学 2024-07-16 Izak Tait , Joshua Bensemann , Ziqi Wang

The goal of building dialogue agents that can converse with humans naturally has been a long-standing dream of researchers since the early days of artificial intelligence. The well-known Turing Test proposed to judge the ultimate validity…

人工智能 · 计算机科学 2022-12-13 Tom Young

This paper investigates the empathetic responding capabilities of ChatGPT, particularly its latest iteration, GPT-4, in comparison to human-generated responses to a wide range of emotional scenarios, both positive and negative. We employ a…

人机交互 · 计算机科学 2024-03-12 Anuradha Welivita , Pearl Pu

As Large Language Models (LLMs) become increasingly integrated into everyday life as general purpose multimodal AI systems, their capabilities to simulate human understanding are under examination. This study investigates LLMs ability to…

计算与语言 · 计算机科学 2025-08-26 Ljubisa Bojic , Predrag Kovacevic , Milan Cabarkapa

We subject GPT-4 to a number of rigorous psychometric tests and analyze the results. We find that, compared to the average human, GPT-4 tends to show more honesty and humility, and less machiavellianism and narcissism. It sometimes exhibits…

计算与语言 · 计算机科学 2024-02-06 Adrita Barua , Gary Brase , Ke Dong , Pascal Hitzler , Eugene Vasserman

Socially fluent agentic AI can now participate in online interaction in ways that resemble ordinary human conversation, potentially weakening people's ability to infer who is human from conversational signals alone. We tested this…

人机交互 · 计算机科学 2026-05-25 Lixiang Yan , Yueqiao Jin , Xibin Han , Dragan Gašević

The best improvisational theatre actors can make any scene partner, of any skill level or ability, appear talented and proficient in the art form, and thus "make them shine". To challenge this improvisational paradigm, we built an…

人工智能 · 计算机科学 2017-12-05 Kory Wallace Mathewson , Piotr Mirowski

Autonomous cars are indispensable when humans go further down the hands-free route. Although existing literature highlights that the acceptance of the autonomous car will increase if it drives in a human-like manner, sparse research offers…

人机交互 · 计算机科学 2023-05-25 Zhaoning Li , Qiaoli Jiang , Zhengming Wu , Anqi Liu , Haiyan Wu , Miner Huang , Kai Huang , Yixuan Ku

Reliable human-machine discrimination is becoming increasingly important as large language models and autonomous agents are deployed in online settings. Existing approaches evaluate whether a system can produce behavior or responses…

人工智能 · 计算机科学 2026-05-12 Milena Rmus , Mathew D. Hardy , Thomas L. Griffiths , Mayank Agrawal

Chess engines passed human strength years ago, but they still don't play like humans. A grandmaster under clock pressure blunders in ways a club player on a hot streak never would. Conventional engines capture none of this. This paper…

人工智能 · 计算机科学 2026-03-06 Diego Armando Resendez Prado

A widespread view is that Artificial Intelligence cannot be creative. We tested this assumption by comparing human-generated ideas with those generated by six Generative Artificial Intelligence (GAI) chatbots: $alpa.\!ai$, $Copy.\!ai$,…

人工智能 · 计算机科学 2023-10-18 Jennifer Haase , Paul H. P. Hanel

Current AI systems minimize risk by enforcing ideological neutrality, yet this may introduce automation bias by suppressing cognitive engagement in human decision-making. We conducted randomized trials with 2,500 participants to test…

人机交互 · 计算机科学 2025-08-21 Shiyang Lai , Junsol Kim , Nadav Kunievsky , Yujin Potter , James Evans

A key objective in artificial intelligence (AI) development is to create systems that match or surpass human creativity. Although current AI models perform well across diverse creative tasks, it remains unclear whether these achievements…

人机交互 · 计算机科学 2025-04-01 Man Zhang , Ying Li , Yang Peng , Yijia Sun , Wenxin Guo , Huiqing Hu , Shi Chen , Qingbai Zhao

Inspired by the Turing test, we present a novel methodological framework to assess the extent to which a population of machines mirrors the philosophical views of a population of humans. The framework consists of three steps: (i)…

物理与社会 · 物理学 2025-10-06 Michele Pizzochero , Giorgia Dellaferrera

Large language models (LLMs) like GPT-4 have recently demonstrated impressive capabilities in natural language understanding and generation. However, there is a concern that they can be misused for malicious purposes, such as fraud or…

计算与语言 · 计算机科学 2024-08-13 Hong Wang , Xuan Luo , Weizhi Wang , Xifeng Yan

Research suggests that providing specific and timely feedback to human tutors enhances their performance. However, it presents challenges due to the time-consuming nature of assessing tutor performance by human evaluators. Large language…

计算与语言 · 计算机科学 2023-07-06 Dollaya Hirunyasiri , Danielle R. Thomas , Jionghao Lin , Kenneth R. Koedinger , Vincent Aleven