中文
相关论文

相关论文: ChatGPT Hallucinates when Attributing Answers

200 篇论文

Large language models (LLMs) are increasingly relied upon to solve complex mathematical word problems. However, being susceptible to hallucination, they may generate inaccurate results when presented with unanswerable questions, raising…

计算与语言 · 计算机科学 2024-10-18 Asir Saadat , Tasmia Binte Sogir , Md Taukir Azam Chowdhury , Syem Aziz

The development of large language models (LLMs) such as ChatGPT has brought a lot of attention recently. However, their evaluation in the benchmark academic datasets remains under-explored due to the difficulty of evaluating the generative…

Large language models (LLMs), including ChatGPT, Bard, and Llama, have achieved remarkable successes over the last two years in a range of different applications. In spite of these successes, there exist concerns that limit the wide…

计算与语言 · 计算机科学 2024-01-17 Junliang Luo , Tianyu Li , Di Wu , Michael Jenkin , Steve Liu , Gregory Dudek

This paper investigates the capabilities of ChatGPT as an automated assistant in diverse domains, including scientific writing, mathematics, education, programming, and healthcare. We explore the potential of ChatGPT to enhance…

人机交互 · 计算机科学 2023-06-07 Amos Azaria , Rina Azoulay , Shulamit Reches

Data visualization creators often lack formal training, resulting in a knowledge gap in design practice. Large language models such as ChatGPT, with their vast internet-scale training data, offer transformative potential to address this…

人机交互 · 计算机科学 2025-05-20 Nam Wook Kim , Yongsu Ahn , Grace Myers , Benjamin Bach

We conducted controlled experimental bias audits for four versions of ChatGPT, which we asked to recommend an opening offer in salary negotiations for a new hire. We submitted 98,800 prompts to each version, systematically varying the…

计算机与社会 · 计算机科学 2024-10-10 R. Stuart Geiger , Flynn O'Sullivan , Elsie Wang , Jonathan Lo

This study aimed to determine if ChatGPT's large language models could match the scoring accuracy of human and machine scores from the ASAP competition. The investigation focused on various prediction models, including linear regression,…

计算与语言 · 计算机科学 2024-08-20 Mark D. Shermis

The objective of legal text entailment is to ascertain whether the assertions in a legal query logically follow from the information provided in one or multiple legal articles. ChatGPT, a large language model, is robust in many natural…

计算与语言 · 计算机科学 2024-02-01 Chau Nguyen , Le-Minh Nguyen

ChatGPT has been emerging as a novel information source, and it is likely that the public might seek information from ChatGPT while taking protective actions when facing climate hazards such as floods and hurricanes. The objective of this…

计算机与社会 · 计算机科学 2023-04-18 Xiangpeng Li , Yuqin Jiang , Ali Mostafavi

Large language models (LLMs) have achieved a degree of success in generating coherent and contextually relevant text, yet they remain prone to a significant challenge known as hallucination: producing information that is not substantiated…

计算与语言 · 计算机科学 2024-10-28 Ray Li , Tanishka Bagade , Kevin Martinez , Flora Yasmin , Grant Ayala , Michael Lam , Kevin Zhu

Since ChatGPT offers detailed responses without justifications, and erroneous facts even for popular persons, events and places, in this paper we present a novel pipeline that retrieves the response of ChatGPT in RDF and tries to validate…

数据库 · 计算机科学 2023-11-20 Michalis Mountantonakis , Yannis Tzitzikas

Recent progress in natural language processing (NLP) owes much to remarkable advances in large language models (LLMs). Nevertheless, LLMs frequently "hallucinate," resulting in non-factual outputs. Our carefully-designed human evaluation…

计算与语言 · 计算机科学 2024-03-22 Jian Guan , Jesse Dodge , David Wadden , Minlie Huang , Hao Peng

Artificial intelligence is gaining traction in more ways than ever before. The popularity of language models and AI-based businesses has soared since ChatGPT was made available to the general public via OpenAI. It is becoming increasingly…

计算机与社会 · 计算机科学 2023-07-31 Prabin Sharma , Kisan Thapa , Dikshya Thapa , Prastab Dhakal , Mala Deep Upadhaya , Santosh Adhikari , Salik Ram Khanal

ChatGPT has demonstrated exceptional proficiency in natural language conversation, e.g., it can answer a wide range of questions while no previous large language models can. Thus, we would like to push its limit and explore its ability to…

计算与语言 · 计算机科学 2023-02-07 Ruibo Tu , Chao Ma , Cheng Zhang

Large-scale language models, like ChatGPT, have garnered significant media attention and stunned the public with their remarkable capacity for generating coherent text from short natural language prompts. In this paper, we aim to conduct a…

计算与语言 · 计算机科学 2024-12-11 Dongqi Liu , Vera Demberg

The release of the large language model based chatbot ChatGPT in November 2022 has brought considerable attention to the subject of artificial intelligence, not only in the public. From the perspective of higher education, ChatGPT…

计算机与社会 · 计算机科学 2023-11-29 Nicolas Schwenke , Heinrich Söbke , Eckhard Kraft

The goal of temporal relation extraction is to infer the temporal relation between two events in the document. Supervised models are dominant in this task. In this work, we investigate ChatGPT's ability on zero-shot temporal relation…

计算与语言 · 计算机科学 2023-04-13 Chenhan Yuan , Qianqian Xie , Sophia Ananiadou

Large language models, like ChatGPT, have shown remarkable capability in many downstream tasks, yet their ability to understand discourse structures of dialogues remains less explored, where it requires higher level capabilities of…

计算与语言 · 计算机科学 2024-03-06 Yaxin Fan , Feng Jiang , Peifeng Li , Haizhou Li

New chat AI applications like ChatGPT offer an advanced understanding of question context and memory across multi-step tasks, such that experiments can test its deductive reasoning. This paper proposes a multi-role and multi-step challenge,…

人工智能 · 计算机科学 2023-01-05 David Noever , Forrest McKee

Large Language Models (LLMs) are increasingly used in software security, but their trustworthiness in generating accurate vulnerability advisories remains uncertain. This study investigates the ability of ChatGPT to (1) generate plausible…