中文
相关论文

相关论文: Interpreting Negation in GPT-2: Layer- and Head-Le…

200 篇论文

Prior interpretability research studying narrow distributions has preliminarily identified self-repair, a phenomena where if components in large language models are ablated, later components will change their behavior to compensate. Our…

机器学习 · 计算机科学 2025-04-15 Cody Rushing , Neel Nanda

Despite the Graph Neural Networks' (GNNs) proficiency in analyzing graph data, achieving high-accuracy and interpretable predictions remains challenging. Existing GNN interpreters typically provide post-hoc explanations disjointed from…

机器学习 · 计算机科学 2024-07-26 Zhenhua Huang , Kunhao Li , Shaojie Wang , Zhaohong Jia , Wentao Zhu , Sharad Mehrotra

Sentiment analysis focuses on identifying the emotional polarity expressed in textual data, typically categorized as positive, negative, or neutral. Hate speech detection, on the other hand, aims to recognize content that incites violence,…

计算与语言 · 计算机科学 2026-01-07 Meysam Shirdel Bilehsavar , Negin Mahmoudi , Mohammad Jalili Torkamani , Kiana Kiashemshaki

Large Language Models (LLMs) are known to exhibit social, demographic, and gender biases, often as a consequence of the data on which they are trained. In this work, we adopt a mechanistic interpretability approach to analyze how such…

计算与语言 · 计算机科学 2025-06-09 Bhavik Chandna , Zubair Bashir , Procheta Sen

Large Language Models such as GPTs (Generative Pre-trained Transformers) exhibit remarkable capabilities across a broad spectrum of applications. Nevertheless, due to their intrinsic complexity, these models present substantial challenges…

机器学习 · 计算机科学 2024-10-17 Ashkan Golgoon , Khashayar Filom , Arjun Ravi Kannan

Natural language explanation (NLE) models aim at explaining the decision-making process of a black box system via generating natural language sentences which are human-friendly, high-level and fine-grained. Current NLE models explain the…

计算机视觉与模式识别 · 计算机科学 2022-03-11 Fawaz Sammani , Tanmoy Mukherjee , Nikos Deligiannis

We investigate the internal structure of language model computations using causal analysis and demonstrate two motifs: (1) a form of adaptive computation where ablations of one attention layer of a language model cause another layer to…

机器学习 · 计算机科学 2023-08-01 Thomas McGrath , Matthew Rahtz , Janos Kramar , Vladimir Mikulik , Shane Legg

Neural conditional language generation models achieve the state-of-the-art in Neural Machine Translation (NMT) but are highly dependent on the quality of parallel training dataset. When trained on low-quality datasets, these models are…

计算与语言 · 计算机科学 2023-06-16 Joël Tang , Marina Fomicheva , Lucia Specia

We explore the topology of representation manifolds arising in autoregressive neural language models trained on raw text data. In order to study their properties, we introduce tools from computational algebraic topology, which we use as a…

计算与语言 · 计算机科学 2024-06-11 Stephen Fitz , Peter Romero , Jiyan Jonas Schneider

According to the embodied cognition perspective, linguistic negation may block the motor simulations induced by language processing. Transcranial magnetic stimulation (TMS) was applied to the left primary motor cortex (hand area) of…

神经元与认知 · 定量生物学 2021-06-09 Giorgio Papitto , Luisa Lugli , Anna M. Borghi , Antonello Pellicano , Ferdinand Binkofski

Mechanistic interpretability seeks to reverse engineer a trained neural network by identifying the minimal subset of internal components. We perform a mechanistic interpretability analysis of the Particle Transformer architecture, trained…

高能物理 - 唯象学 · 物理学 2026-05-12 Saurabh Rai , Sanmay Ganguly

Common methods for interpreting neural models in natural language processing typically examine either their structure or their behavior, but not both. We propose a methodology grounded in the theory of causal mediation analysis for…

Negation instructions such as 'do not mention $X$' can paradoxically increase the accessibility of $X$ in human thought, a phenomenon known as ironic rebound. Large language models (LLMs) face the same challenge: suppressing a concept…

计算与语言 · 计算机科学 2025-11-18 Logan Mann , Nayan Saxena , Sarah Tandon , Chenhao Sun , Savar Toteja , Kevin Zhu

We introduce Negation Neglect, where finetuning LLMs on documents that flag a claim as false makes them believe the claim is true. For example, models are finetuned on documents that convey "Ed Sheeran won the 100m gold at the 2024…

计算与语言 · 计算机科学 2026-05-14 Harry Mayne , Lev McKinney , Jan Dubiński , Adam Karvonen , James Chua , Owain Evans

Previous works of negation understanding mainly focus on negation cue detection and scope resolution, without identifying negation subject which is also significant to the downstream tasks. In this paper, we propose a new negation triplet…

计算与语言 · 计算机科学 2024-04-16 Yuchen Shi , Deqing Yang , Jingping Liu , Yanghua Xiao , Zongyu Wang , Huimin Xu

We propose a framework to model an operational conversational negation by applying worldly context (prior knowledge) to logical negation in compositional distributional semantics. Given a word, our framework can create its negation that is…

计算与语言 · 计算机科学 2021-05-17 Benjamin Rodatz , Razin A. Shaikh , Lia Yeh

The logical negation property (LNP), which implies generating different predictions for semantically opposite inputs, is an important property that a trustworthy language model must satisfy. However, much recent evidence shows that…

计算与语言 · 计算机科学 2022-08-12 Myeongjun Jang , Frank Mtumbuka , Thomas Lukasiewicz

A number of recent benchmarks seek to assess how well models handle natural language negation. However, these benchmarks lack the controlled example paradigms that would allow us to infer whether a model had learned how negation morphemes…

计算与语言 · 计算机科学 2024-04-19 Jingyuan Selena She , Christopher Potts , Samuel R. Bowman , Atticus Geiger

Safety and controllability are critical for large language models. A central question is whether undesirable behaviors like deception are localized functions that can be removed, or if they are deeply intertwined with a model's core…

计算与语言 · 计算机科学 2025-10-01 Eduard Kapelko

Two fundamental questions in neurolinguistics concerns the brain regions that integrate information beyond the lexical level, and the size of their window of integration. To address these questions we introduce a new approach named…

计算与语言 · 计算机科学 2023-05-24 Alexandre Pasquiou , Yair Lakretz , Bertrand Thirion , Christophe Pallier