English
Related papers

Related papers: How Transformers Reject Wrong Answers: Rotational …

200 papers

Why do language models trained on contradictory data prefer correct answers? In controlled experiments with small transformers (3.5M--86M parameters), we show that this preference tracks the compressibility structure of errors rather than…

Computation and Language · Computer Science 2026-04-07 Konstantin Krestnikov

Factual incorrectness in generated content is one of the primary concerns in ubiquitous deployment of large language models (LLMs). Prior findings suggest LLMs can (sometimes) detect factual incorrectness in their generated content (i.e.,…

Computation and Language · Computer Science 2025-05-28 Hovhannes Tamoyan , Subhabrata Dutta , Iryna Gurevych

Reasoning models are evaluated on single-turn benchmarks but deployed in multi-turn dialogue, where users push back on correct answers. Under sustained adversarial pressure we find a previously undocumented failure mode: the…

Artificial Intelligence · Computer Science 2026-05-29 Yubo Li , Ramayya Krishnan , Rema Padman

Language models are becoming the default interface to factual knowledge, yet they often verify outputs more reliably than they generate them. This generation-verification gap (GV-gap) underlies many recent advances in self-improvement and…

Computation and Language · Computer Science 2026-05-28 Tim R. Davidson , Anja Surina , Caglar Gulcehre

Current alignment evaluation mostly measures whether models encode dangerous concepts and whether they refuse harmful requests. Both miss the layer where alignment often operates: routing from concept detection to behavioral policy. We…

Machine Learning · Computer Science 2026-05-04 Gregory N. Frank

Reasoning models have attracted increasing attention for their ability to tackle complex tasks, embodying the System II (slow thinking) paradigm in contrast to System I (fast, intuitive responses). Yet a key question remains: Does slower…

Artificial Intelligence · Computer Science 2026-04-17 Sitong Fang , Wenjing Cao , Jiahao Li , Xuyao Wang , Juntao Dai , Chi-Min Chan , Sirui Han , Yike Guo , Yaodong Yang , Jiaming Ji

Existing explainability methods for Large Language Models (LLMs) typically treat hidden states as static points in activation space, assuming that correct and incorrect inferences can be separated using representations from an individual…

Computation and Language · Computer Science 2026-03-03 Hamed Damirchi , Ignacio Meza De la Jara , Ehsan Abbasnejad , Afshar Shamsi , Zhen Zhang , Javen Shi

Large Language Models (LLMs) have impressive capabilities, but are prone to outputting falsehoods. Recent work has developed techniques for inferring whether a LLM is telling the truth by training probes on the LLM's internal activations.…

Artificial Intelligence · Computer Science 2024-08-20 Samuel Marks , Max Tegmark

Predicting upcoming events is critical to our ability to interact with our environment. Transformer models, trained on next-word prediction, appear to construct representations of linguistic input that can support diverse downstream tasks.…

Computation and Language · Computer Science 2023-11-10 Eghbal A. Hosseini , Evelina Fedorenko

LLMs reliably correct false claims when presented in isolation, yet when the same claims are embedded in task-oriented requests, they often comply rather than correct. We term this failure mode \emph{correction suppression} and construct a…

Machine Learning · Computer Science 2026-05-11 Zixuan Chen , Hao Lin , Zizhe Chen , Yizhou Tian , Garry Yang , Depeng Wang , Ya Guo , Huijia Zhu , James Cheng

We discover that large language models exhibit \emph{spectral phase transitions} in their hidden activation spaces when engaging in reasoning versus factual recall. Through systematic spectral analysis across \textbf{11 models} spanning…

Machine Learning · Computer Science 2026-04-20 Yi Liu

Model transformations operate on models conforming to precisely defined metamodels. Consequently, it often seems relatively easy to chain them: the output of a transformation may be given as input to a second one if metamodels match.…

Artificial Intelligence · Computer Science 2010-03-04 Raphael Chenouard , Frédéric Jouault

Despite advancements in large language models (LLMs), non-factual responses still persist in fact-seeking question answering. Unlike extensive studies on post-hoc detection of these responses, this work studies non-factuality prediction…

Computation and Language · Computer Science 2025-08-19 Yanling Wang , Haoyang Li , Hao Zou , Jing Zhang , Xinlei He , Qi Li , Ke Xu

Bidirectional Encoder Representations from Transformers (BERT) reach state-of-the-art results in a variety of Natural Language Processing tasks. However, understanding of their internal functioning is still insufficient and unsatisfactory.…

Computation and Language · Computer Science 2019-09-12 Betty van Aken , Benjamin Winter , Alexander Löser , Felix A. Gers

For Large Language Models (LLMs) to be reliably deployed, models must effectively know when not to answer: abstain. Reasoning models, in particular, have gained attention for impressive performance on complex tasks. However, reasoning…

Artificial Intelligence · Computer Science 2026-04-03 Abinitha Gourabathina , Inkit Padhi , Manish Nagireddy , Subhajit Chaudhury , Prasanna Sattigeri

Transformers have significantly advanced the field of natural language processing, but comprehending their internal mechanisms remains a challenge. In this paper, we introduce a novel geometric perspective that elucidates the inner…

Computation and Language · Computer Science 2023-09-20 Raul Molina

Understanding subjectivity demands reasoning skills beyond the realm of common knowledge. It requires a machine learning model to process sentiment and to perform opinion mining. In this work, I've exploited a recently released dataset for…

Computation and Language · Computer Science 2020-10-15 Lukas Muttenthaler

Despite recent success, large neural models often generate factually incorrect text. Compounding this is the lack of a standard automatic evaluation for factuality--it cannot be meaningfully improved if it cannot be measured. Grounded…

Computation and Language · Computer Science 2022-03-30 Peter West , Chris Quirk , Michel Galley , Yejin Choi

Large Language Models (LLMs) exhibit strong conversational abilities but often generate falsehoods. Prior work suggests that the truthfulness of simple propositions can be represented as a single linear direction in a model's internal…

Machine Learning · Computer Science 2025-05-29 Stanley Yu , Vaidehi Bulusu , Oscar Yasunaga , Clayton Lau , Cole Blondin , Sean O'Brien , Kevin Zhu , Vasu Sharma

Language model representations often contain linear directions that correspond to high-level concepts. Here, we study the dynamics of these representations: how representations evolve along these dimensions within the context of (simulated)…

Computation and Language · Computer Science 2026-02-04 Andrew Kyle Lampinen , Yuxuan Li , Eghbal Hosseini , Sangnie Bhardwaj , Murray Shanahan
‹ Prev 1 2 3 10 Next ›