English
Related papers

Related papers: Truth Neurons

200 papers

Transformers have been shown to be able to perform deductive reasoning on a logical rulebase containing rules and statements written in English natural language. While the progress is promising, it is currently unclear if these models…

Computation and Language · Computer Science 2022-11-09 Soumya Sanyal , Zeyi Liao , Xiang Ren

Large Language Models (LLMs) encode vast world knowledge across multiple languages, yet their internal beliefs are often unevenly distributed across linguistic spaces. When external evidence contradicts these language-dependent memories,…

Computation and Language · Computer Science 2026-01-13 Jiaqi Zhao , Qiang Huang , Haodong Chen , Xiaoxing You , Jun Yu

Evaluating language models and AI agents remains fundamentally challenging because static benchmarks fail to capture real-world uncertainty, distribution shift, and the gap between isolated task accuracy and human-aligned decision-making…

Artificial Intelligence · Computer Science 2026-01-27 Shirin Shahabi , Spencer Graham , Haruna Isah

Language Models (LMs) may acquire harmful knowledge, and yet feign ignorance of these topics when under audit. Inspired by the recent discovery of deception-related behaviour patterns in LMs, we aim to train classifiers that detect when a…

Computation and Language · Computer Science 2026-03-24 Dhananjay Ashok , Ruth-Ann Armstrong , Jonathan May

Coherent discourse is distinguished from a mere collection of utterances by the satisfaction of a diverse set of constraints, for example choice of expression, logical relation between denoted events, and implicit compatibility with…

Computation and Language · Computer Science 2021-05-11 Anne Beyer , Sharid Loáiciga , David Schlangen

Large Language Models (LLMs) achieve remarkable performance across various tasks, but their tendency to produce hallucinations limits reliable adoption. Benchmarks such as TruthfulQA have been developed to measure truthfulness, yet they are…

Computation and Language · Computer Science 2025-09-09 Lorenzo Alfred Nery , Ronald Dawson Catignas , Thomas James Tiam-Lee

Transformer language models have received widespread public attention, yet their generated text is often surprising even to NLP researchers. In this survey, we discuss over 250 recent studies of English language model behavior before…

Computation and Language · Computer Science 2023-08-29 Tyler A. Chang , Benjamin K. Bergen

Instruction-following language models often show undesirable biases. These undesirable biases may be accelerated in the real-world usage of language models, where a wide range of instructions is used through zero-shot example prompting. To…

Artificial Intelligence · Computer Science 2024-06-06 Nakyeong Yang , Taegwan Kang , Jungkyu Choi , Honglak Lee , Kyomin Jung

Formal verification provides critical security assurances for neural networks, yet its practical application suffers from the long verification time. This work introduces a novel method for training verification-friendly neural networks,…

Machine Learning · Computer Science 2024-12-31 Zongxin Liu , Zhe Zhao , Fu Song , Jun Sun , Pengfei Yang , Xiaowei Huang , Lijun Zhang

LLMs frequently generate fictitious yet convincing citations, often expressing high confidence even when the underlying reference is wrong. We study this failure across 9 models and 108{,}000 generated references, and find that author names…

Computation and Language · Computer Science 2026-04-22 Yuefei Chen , Yihao Quan , Xiaodong Lin , Ruixiang Tang

We investigate how low-quality AI advisors, lacking quality disclosures, can help spread text-based lies while seeming to help people detect lies. Participants in our experiment discern truth from lies by evaluating transcripts from a game…

Computation and Language · Computer Science 2025-02-04 Haimanti Bhattacharya , Subhasish Dugar , Sanchaita Hazra , Bodhisattwa Prasad Majumder

Large Language Models (LLMs) store an extensive amount of factual knowledge obtained from vast collections of text. To effectively utilize these models for downstream tasks, it is crucial to have reliable methods for measuring their…

Computation and Language · Computer Science 2023-06-13 Pouya Pezeshkpour

The capabilities of large language models (LLMs) have sparked debate over whether such systems just learn an enormous collection of superficial statistics or a set of more coherent and grounded representations that reflect the real world.…

Machine Learning · Computer Science 2024-03-05 Wes Gurnee , Max Tegmark

A critical step to building trustworthy deep neural networks is trust quantification, where we ask the question: How much can we trust a deep neural network? In this study, we take a step towards simple, interpretable metrics for trust…

Machine Learning · Computer Science 2021-04-06 Alexander Wong , Xiao Yu Wang , Andrew Hryniowski

Large language models (LLMs) exhibit remarkable similarity to neural activity in the human language network. However, the key properties of language shaping brain-like representations, and their evolution during training as a function of…

Computation and Language · Computer Science 2025-09-23 Badr AlKhamissi , Greta Tuckute , Yingtian Tang , Taha Binhuraib , Antoine Bosselut , Martin Schrimpf

While improving neural dialogue agents' factual accuracy is the object of much research, another important aspect of communication, less studied in the setting of neural dialogue, is transparency about ignorance. In this work, we analyze to…

Computation and Language · Computer Science 2022-06-28 Sabrina J. Mielke , Arthur Szlam , Emily Dinan , Y-Lan Boureau

Large Language Models (LLMs) exhibit strong conversational abilities but often generate falsehoods. Prior work suggests that the truthfulness of simple propositions can be represented as a single linear direction in a model's internal…

Machine Learning · Computer Science 2025-05-29 Stanley Yu , Vaidehi Bulusu , Oscar Yasunaga , Clayton Lau , Cole Blondin , Sean O'Brien , Kevin Zhu , Vasu Sharma

Rapid integration of large language models (LLMs) into societal applications has intensified concerns about their alignment with universal ethical principles, as their internal value representations remain opaque despite behavioral…

Computation and Language · Computer Science 2025-05-26 Yi Su , Jiayi Zhang , Shu Yang , Xinhai Wang , Lijie Hu , Di Wang

It has been an open question in deep learning if fault-tolerant computation is possible: can arbitrarily reliable computation be achieved using only unreliable neurons? In the grid cells of the mammalian cortex, analog error correction…

Machine Learning · Computer Science 2025-03-26 Alexander Zlokapa , Andrew K. Tan , John M. Martyn , Ila R. Fiete , Max Tegmark , Isaac L. Chuang

The spread of digital disinformation (aka "fake news") is arguably one of the most significant threats on the Internet which can cause individual and societal harm of large scales. The susceptibility to fake news attacks hinges on whether…

Computation and Language · Computer Science 2022-07-19 Cagri Arisoy , Anuradha Mandal , Nitesh Saxena
‹ Prev 1 8 9 10 Next ›