中文
相关论文

相关论文: True Lies

200 篇论文

Large language models (LLMs) are trained on vast amounts of text from the internet, which contains both factual and misleading information about the world. While unintuitive from a classic view of LMs, recent work has shown that the truth…

计算与语言 · 计算机科学 2024-02-07 Nitish Joshi , Javier Rando , Abulhair Saparov , Najoung Kim , He He

In standard epistemic logic, knowing that p is the same as knowing that p is true, but it does not say anything about understanding p or knowing its meaning. In this paper, we present a conservative extension of Public Announcement Logic…

计算机科学中的逻辑 · 计算机科学 2019-07-23 Malvin Gattinger , Yanjing Wang

Large language models (LLMs) can "lie", which we define as outputting false statements despite "knowing" the truth in a demonstrable sense. LLMs might "lie", for example, when instructed to output misinformation. Here, we develop a simple…

Common sense suggests that when individuals explain why they believe something, we can arrive at more accurate conclusions than when they simply state what they believe. Yet, there is no known mechanism that provides incentives to elicit…

计算机科学与博弈论 · 计算机科学 2025-02-20 Siddarth Srinivasan , Ezra Karger , Michiel Bakker , Yiling Chen

Neural language models (LMs) can be used to evaluate the truth of factual statements in two ways: they can be either queried for statement probabilities, or probed for internal representations of truthfulness. Past work has found that these…

计算与语言 · 计算机科学 2023-12-08 Kevin Liu , Stephen Casper , Dylan Hadfield-Menell , Jacob Andreas

Deceptive agents are a challenge for the safety, trustworthiness, and cooperation of AI systems. We focus on the problem that agents might deceive in order to achieve their goals (for instance, in our experiments with language models, the…

人工智能 · 计算机科学 2023-12-05 Francis Rhys Ward , Francesco Belardinelli , Francesca Toni , Tom Everitt

Traditionally, an agent's beliefs would come from what the agent can see, hear, or sense. In the modern world, beliefs are often based on the data available to the agents. In this work, we investigate a dynamic logic of such beliefs that…

计算机科学中的逻辑 · 计算机科学 2025-11-04 Junli Jiang , Pavel Naumov , Wenxuan Zhang

We study an extension of the voter model in which each agent is endowed with an innate preference for one of two states that we term as "truth" or "falsehood". Due to interactions with neighbors, an agent that innately prefers truth can be…

物理与社会 · 物理学 2011-05-03 Naoki Masuda , S. Redner

I outline a new theory of truth that resolves the classical and constructive versions of the liar paradox. The theory features a provably consistent axiomatization of a global self-applicative truth predicate. Truth is defined using…

逻辑 · 数学 2025-07-14 Nik Weaver

Agents' judgment depends on perception and previous knowledge. Assuming that previous knowledge depends on perception, we can say that judgment depends on perception. So, if judgment depends on perception, can agents judge that they have…

神经元与认知 · 定量生物学 2012-02-21 Ahmed M. Mahran

The speech-to-song illusion is a robust psychological phenomenon whereby a spoken sentence sounds increasingly more musical as it is repeated. Despite decades of research, a complete formal account of this transformation is still lacking,…

神经元与认知 · 定量生物学 2024-02-13 Raja Marjieh , Pol van Rijn , Ilia Sucholutsky , Harin Lee , Thomas L. Griffiths , Nori Jacoby

We study the structure of the partial order induced by the definability relation on definitions of truth for the language of arithmetic. Formally, a definition of truth is any sentence $\alpha$ which extends a weak arithmetical theory…

逻辑 · 数学 2023-11-23 Piotr Gruza , Mateusz Łełyk

We produce a connected real Lie group that, as a first order structure in the group language, interprets the real field expanded with a predicate for the integers. Moreover, the domain of our interpretation is definable in the group.

逻辑 · 数学 2021-08-20 Annalisa Conversano , Marcello Mamino

In many contexts, lying -- the use of verbal falsehoods to deceive -- is harmful. While lying has traditionally been a human affair, AI systems that make sophisticated verbal statements are becoming increasingly prevalent. This raises the…

计算机与社会 · 计算机科学 2021-10-14 Owain Evans , Owen Cotton-Barratt , Lukas Finnveden , Adam Bales , Avital Balwit , Peter Wills , Luca Righetti , William Saunders

A definition of what counts as an explanation of mathematical statement, and when one explanation is better than another, is given. Since all mathematical facts must be true in all causal models, and hence known by an agent, mathematical…

人工智能 · 计算机科学 2024-02-16 Joseph Y. Halpern

Truthfulness (adherence to factual accuracy) and utility (satisfying human needs and instructions) are both fundamental aspects of Large Language Models, yet these goals often conflict (e.g., sell a car with known flaws), which makes it…

人工智能 · 计算机科学 2025-04-29 Zhe Su , Xuhui Zhou , Sanketh Rangreji , Anubha Kabra , Julia Mendelsohn , Faeze Brahman , Maarten Sap

Defining various dishonest notions in a formal way is a key step to enable intelligent agents to act in untrustworthy environments. This review evaluates the literature for this topic by looking at formal definitions based on modal logic as…

计算机科学中的逻辑 · 计算机科学 2016-12-30 Toni Heidenreich

This paper studies knowledge representation in multi-agent environment. We investigate technique for computation truth-values of statements based at a new temporal, agent's-knowledge logic TL. A logical language, mathematical symbolic…

计算机科学中的逻辑 · 计算机科学 2014-06-23 Maybin Muyeba , Vladimir Rybakov

Automated fact checking systems have been proposed that quickly provide veracity prediction at scale to mitigate the negative influence of fake news on people and on public opinion. However, most studies focus on veracity classifiers of…

计算与语言 · 计算机科学 2022-06-15 Shih-Chieh Dai , Yi-Li Hsu , Aiping Xiong , Lun-Wei Ku

Dynamic epistemic logics which model abilities of agents to make various announcements and influence each other's knowledge have been studied extensively in recent years. Two notable examples of such logics are Group Announcement Logic and…

计算机科学中的逻辑 · 计算机科学 2018-10-08 Rustam Galimullin , Natasha Alechina