中文
相关论文

相关论文: Truth Neurons

200 篇论文

Transformers have been shown to be able to perform deductive reasoning on a logical rulebase containing rules and statements written in English natural language. While the progress is promising, it is currently unclear if these models…

计算与语言 · 计算机科学 2022-11-09 Soumya Sanyal , Zeyi Liao , Xiang Ren

Large Language Models (LLMs) encode vast world knowledge across multiple languages, yet their internal beliefs are often unevenly distributed across linguistic spaces. When external evidence contradicts these language-dependent memories,…

计算与语言 · 计算机科学 2026-01-13 Jiaqi Zhao , Qiang Huang , Haodong Chen , Xiaoxing You , Jun Yu

Evaluating language models and AI agents remains fundamentally challenging because static benchmarks fail to capture real-world uncertainty, distribution shift, and the gap between isolated task accuracy and human-aligned decision-making…

人工智能 · 计算机科学 2026-01-27 Shirin Shahabi , Spencer Graham , Haruna Isah

Language Models (LMs) may acquire harmful knowledge, and yet feign ignorance of these topics when under audit. Inspired by the recent discovery of deception-related behaviour patterns in LMs, we aim to train classifiers that detect when a…

计算与语言 · 计算机科学 2026-03-24 Dhananjay Ashok , Ruth-Ann Armstrong , Jonathan May

Coherent discourse is distinguished from a mere collection of utterances by the satisfaction of a diverse set of constraints, for example choice of expression, logical relation between denoted events, and implicit compatibility with…

计算与语言 · 计算机科学 2021-05-11 Anne Beyer , Sharid Loáiciga , David Schlangen

Large Language Models (LLMs) achieve remarkable performance across various tasks, but their tendency to produce hallucinations limits reliable adoption. Benchmarks such as TruthfulQA have been developed to measure truthfulness, yet they are…

计算与语言 · 计算机科学 2025-09-09 Lorenzo Alfred Nery , Ronald Dawson Catignas , Thomas James Tiam-Lee

Transformer language models have received widespread public attention, yet their generated text is often surprising even to NLP researchers. In this survey, we discuss over 250 recent studies of English language model behavior before…

计算与语言 · 计算机科学 2023-08-29 Tyler A. Chang , Benjamin K. Bergen

Instruction-following language models often show undesirable biases. These undesirable biases may be accelerated in the real-world usage of language models, where a wide range of instructions is used through zero-shot example prompting. To…

人工智能 · 计算机科学 2024-06-06 Nakyeong Yang , Taegwan Kang , Jungkyu Choi , Honglak Lee , Kyomin Jung

Formal verification provides critical security assurances for neural networks, yet its practical application suffers from the long verification time. This work introduces a novel method for training verification-friendly neural networks,…

机器学习 · 计算机科学 2024-12-31 Zongxin Liu , Zhe Zhao , Fu Song , Jun Sun , Pengfei Yang , Xiaowei Huang , Lijun Zhang

LLMs frequently generate fictitious yet convincing citations, often expressing high confidence even when the underlying reference is wrong. We study this failure across 9 models and 108{,}000 generated references, and find that author names…

计算与语言 · 计算机科学 2026-04-22 Yuefei Chen , Yihao Quan , Xiaodong Lin , Ruixiang Tang

We investigate how low-quality AI advisors, lacking quality disclosures, can help spread text-based lies while seeming to help people detect lies. Participants in our experiment discern truth from lies by evaluating transcripts from a game…

计算与语言 · 计算机科学 2025-02-04 Haimanti Bhattacharya , Subhasish Dugar , Sanchaita Hazra , Bodhisattwa Prasad Majumder

Large Language Models (LLMs) store an extensive amount of factual knowledge obtained from vast collections of text. To effectively utilize these models for downstream tasks, it is crucial to have reliable methods for measuring their…

计算与语言 · 计算机科学 2023-06-13 Pouya Pezeshkpour

The capabilities of large language models (LLMs) have sparked debate over whether such systems just learn an enormous collection of superficial statistics or a set of more coherent and grounded representations that reflect the real world.…

机器学习 · 计算机科学 2024-03-05 Wes Gurnee , Max Tegmark

A critical step to building trustworthy deep neural networks is trust quantification, where we ask the question: How much can we trust a deep neural network? In this study, we take a step towards simple, interpretable metrics for trust…

机器学习 · 计算机科学 2021-04-06 Alexander Wong , Xiao Yu Wang , Andrew Hryniowski

Large language models (LLMs) exhibit remarkable similarity to neural activity in the human language network. However, the key properties of language shaping brain-like representations, and their evolution during training as a function of…

计算与语言 · 计算机科学 2025-09-23 Badr AlKhamissi , Greta Tuckute , Yingtian Tang , Taha Binhuraib , Antoine Bosselut , Martin Schrimpf

While improving neural dialogue agents' factual accuracy is the object of much research, another important aspect of communication, less studied in the setting of neural dialogue, is transparency about ignorance. In this work, we analyze to…

计算与语言 · 计算机科学 2022-06-28 Sabrina J. Mielke , Arthur Szlam , Emily Dinan , Y-Lan Boureau

Large Language Models (LLMs) exhibit strong conversational abilities but often generate falsehoods. Prior work suggests that the truthfulness of simple propositions can be represented as a single linear direction in a model's internal…

机器学习 · 计算机科学 2025-05-29 Stanley Yu , Vaidehi Bulusu , Oscar Yasunaga , Clayton Lau , Cole Blondin , Sean O'Brien , Kevin Zhu , Vasu Sharma

Rapid integration of large language models (LLMs) into societal applications has intensified concerns about their alignment with universal ethical principles, as their internal value representations remain opaque despite behavioral…

计算与语言 · 计算机科学 2025-05-26 Yi Su , Jiayi Zhang , Shu Yang , Xinhai Wang , Lijie Hu , Di Wang

It has been an open question in deep learning if fault-tolerant computation is possible: can arbitrarily reliable computation be achieved using only unreliable neurons? In the grid cells of the mammalian cortex, analog error correction…

机器学习 · 计算机科学 2025-03-26 Alexander Zlokapa , Andrew K. Tan , John M. Martyn , Ila R. Fiete , Max Tegmark , Isaac L. Chuang

The spread of digital disinformation (aka "fake news") is arguably one of the most significant threats on the Internet which can cause individual and societal harm of large scales. The susceptibility to fake news attacks hinges on whether…

计算与语言 · 计算机科学 2022-07-19 Cagri Arisoy , Anuradha Mandal , Nitesh Saxena
‹ 上一页 1 8 9 10 下一页 ›