中文
相关论文

相关论文: Latent Introspection: Models Can Detect Prior Conc…

200 篇论文

Systems of interacting continuous-time Markov chains are a powerful model class, but inference is typically intractable in high dimensional settings. Auxiliary information, such as noisy observations, is typically only available at discrete…

机器学习 · 统计学 2026-04-21 Giosue Migliorini , Padhraic Smyth

Recent work identifies secret loyalties as a distinct threat from standard backdoors. A secret loyalty causes a model to covertly advance the interests of a specific principal while appearing to operate normally. We construct the first…

密码学与安全 · 计算机科学 2026-05-13 Alfie Lamerton , Fabien Roger

The advancement of multimodal large language models (MLLMs) has enabled impressive perception capabilities. However, their reasoning process often remains a "fast thinking" paradigm, reliant on end-to-end generation or explicit,…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Yiming Zhang , Qiangyu Yan , Borui Jiang , Kai Han

Attention mechanisms have recently boosted performance on a range of NLP tasks. Because attention layers explicitly weight input components' representations, it is also often assumed that attention can be used to identify information that…

计算与语言 · 计算机科学 2019-06-11 Sofia Serrano , Noah A. Smith

The opaque nature of Large Language Models (LLMs) has led to significant research efforts aimed at enhancing their interpretability, primarily through post-hoc methods. More recent in-hoc approaches, such as Concept Bottleneck Models…

机器学习 · 计算机科学 2025-02-20 Or Raphael Bidusa , Shaul Markovitch

The growing sophistication of synthetic image and deepfake generation models has turned source attribution and authenticity verification into a critical challenge for modern computer vision systems. Recent studies suggest that diffusion…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Claudio Giusti , Luca Guarnera , Sebastiano Battiato

As large language models become increasingly capable, there is growing concern that they may develop reasoning processes that are encoded or hidden from human oversight. To investigate whether current interpretability techniques can…

人工智能 · 计算机科学 2025-12-09 Ching Fang , Samuel Marks

Reasoning models excel at complex tasks such as coding and mathematics, yet their inference is often slow and token-inefficient. To improve the inference efficiency, post-training quantization (PTQ) usually comes with the cost of large…

机器学习 · 计算机科学 2026-01-22 Keyu Lv , Manyi Zhang , Xiaobo Xia , Jingchen Ni , Shannan Yan , Xianzhi Yu , Lu Hou , Chun Yuan , Haoli Bai

The proliferation of text-to-image diffusion models (T2I DMs) has led to an increased presence of AI-generated images in daily life. However, biased T2I models can generate content with specific tendencies, potentially influencing people's…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Huayang Huang , Xiangye Jin , Jiaxu Miao , Yu Wu

Many datasets have been shown to contain incidental correlations created by idiosyncrasies in the data collection process. For example, sentence entailment datasets can have spurious word-class correlations if nearly all contradiction…

机器学习 · 计算机科学 2020-11-10 Christopher Clark , Mark Yatskar , Luke Zettlemoyer

Recent advances in long-context language models (LCLMs), designed to handle extremely long contexts, primarily focus on utilizing external contextual information, often leaving the influence of language models' parametric knowledge…

计算与语言 · 计算机科学 2026-02-09 Yu Fu , Haz Sameen Shahgir , Hui Liu , Xianfeng Tang , Qi He , Yue Dong

Implicit biases refer to automatic mental processes that shape perceptions, judgments, and behaviors. Previous research on "implicit bias" in LLMs focused primarily on outputs rather than the processes underlying the outputs. We present the…

计算机与社会 · 计算机科学 2026-04-07 Messi H. J. Lee , Calvin K. Lai

Recent reasoning language models, particularly those that employ long latent chains of thought, achieve strong performance on complex agentic tasks. However, as these models operate over increasingly long time horizons, their internal…

机器学习 · 计算机科学 2026-05-27 Hans Peter Lyngsøe Raaschou-Jensen , Constanza Fierro , Anders Søgaard

We show that continual pretraining on plausible misinformation can overwrite specific factual knowledge in large language models without degrading overall performance. Unlike prior poisoning work under static pretraining, we study repeated…

机器学习 · 计算机科学 2026-02-09 Svetlana Churina , Niranjan Chebrolu , Kokil Jaidka

This paper introduces a novel self-consciousness defense mechanism for Large Language Models (LLMs) to combat prompt injection attacks. Unlike traditional approaches that rely on external classifiers, our method leverages the LLM's inherent…

人工智能 · 计算机科学 2025-10-03 Boshi Huang , Fabio Nonato de Paula

This research aims to unravel how large language models (LLMs) iteratively refine token predictions through internal processing. We utilized a logit lens technique to analyze the model's token predictions derived from intermediate…

计算与语言 · 计算机科学 2025-06-10 Jaturong Kongmanee

Contextual hallucinations -- statements unsupported by given context -- remain a significant challenge in AI. We demonstrate a practical interpretability insight: a generator-agnostic observer model detects hallucinations via a single…

机器学习 · 计算机科学 2025-08-01 Charles O'Neill , Slava Chalnev , Chi Chi Zhao , Max Kirkby , Mudith Jayasekara

Concept bottleneck models (CBMs) ensure interpretability by decomposing predictions into human interpretable concepts. Yet the annotations used for training CBMs that enable this transparency are often noisy, and the impact of such…

机器学习 · 计算机科学 2026-02-02 Seonghwan Park , Jueun Mun , Donghyun Oh , Namhoon Lee

Perceptual estimates exhibit a reversal in bias depending on uncertainty: they shift toward prior expectations under high stimulus noise, but away from them when sensory noise dominates. The normative framework of a Bayesian observer model…

神经元与认知 · 定量生物学 2025-10-16 Hyun-Jun Jeon , Hansol Choi , Oh-Sang Kwon

We study the evolution of latent space in fine-tuned NLP models. Different from the commonly used probing-framework, we opt for an unsupervised method to analyze representations. More specifically, we discover latent concepts in the…

计算与语言 · 计算机科学 2022-10-25 Nadir Durrani , Hassan Sajjad , Fahim Dalvi , Firoj Alam