中文
相关论文

相关论文: Emergent Introspection in AI is Content-Agnostic

200 篇论文

We uncover a latent capacity for introspection in a Qwen 32B model, demonstrating that the model can detect when concepts have been injected into its earlier context and identify which concept was injected. While the model denies injection…

人工智能 · 计算机科学 2026-02-27 Theia Pearson-Vogel , Martin Vanek , Raymond Douglas , Jan Kulveit

We investigate whether large language models can introspect on their internal states. It is difficult to answer this question through conversation alone, as genuine introspection cannot be distinguished from confabulations. Here, we address…

计算与语言 · 计算机科学 2026-01-06 Jack Lindsey

A hallmark of human intelligence is Introspection-the ability to assess and reason about one's own cognitive processes. Introspection has emerged as a promising but contested capability in large language models (LLMs). However, current…

人工智能 · 计算机科学 2026-03-24 Atharv Naphade , Samarth Bhargav , Sean Lim , Mcnair Shah

Can large language models introspect, that is, accurately detect perturbations to their own internal states? We systematically investigate this question using activation steering in Meta-Llama-3.1-8B-Instruct. First, we show that the binary…

人工智能 · 计算机科学 2026-03-03 Ely Hahami , Ishaan Sinha , Lavik Jain , Josh Kaplan , Jon Hahami

Whether AI models can introspect is an increasingly important practical question. But there is no consensus on how introspection is to be defined. Beginning from a recently proposed ''lightweight'' definition, we argue instead for a thicker…

人工智能 · 计算机科学 2025-08-21 Siyuan Song , Harvey Lederman , Jennifer Hu , Kyle Mahowald

Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introspective awareness." We investigate the mechanisms underlying…

机器学习 · 计算机科学 2026-05-18 Uzay Macar , Li Yang , Atticus Wang , Peter Wallich , Emmanuel Ameisen , Jack Lindsey

There has been recent interest in whether large language models (LLMs) can introspect about their own internal states. Such abilities would make LLMs more interpretable, and also validate the use of standard introspective methods in…

计算与语言 · 计算机科学 2025-09-25 Siyuan Song , Jennifer Hu , Kyle Mahowald

Humans acquire knowledge by observing the external world, but also by introspection. Introspection gives a person privileged access to their current state of mind (e.g., thoughts and feelings) that is not accessible to external observers.…

计算与语言 · 计算机科学 2024-10-18 Felix J Binder , James Chua , Tomek Korbak , Henry Sleight , John Hughes , Robert Long , Ethan Perez , Miles Turpin , Owain Evans

In decision making tasks under uncertainty, humans display characteristic biases in seeking, integrating, and acting upon information relevant to the task. Here, we reexamine data from previous carefully designed experiments, collected at…

人工智能 · 计算机科学 2021-02-05 Soumya Chatterjee , Pradeep Shenoy

Can large language models detect and report their own internal states? A number of studies have argued that the answer to this question is yes. We argue, based on lessons from human metacognition research, that this conclusion may be…

人工智能 · 计算机科学 2026-05-27 Shashwat Singh , Tal Linzen , Shauli Ravfogel

Self-consciousness, the introspection of one's existence and thoughts, represents a high-level cognitive process. As language models advance at an unprecedented pace, a critical question arises: Are these models becoming self-conscious?…

计算与语言 · 计算机科学 2024-10-25 Sirui Chen , Shu Yu , Shengjie Zhao , Chaochao Lu

Do reasoning models have "Aha!" moments? Prior work suggests that models like DeepSeek-R1-Zero undergo sudden mid-trace realizations that lead to accurate outputs, implying an intrinsic capacity for self-correction. Yet, it remains unclear…

人工智能 · 计算机科学 2026-04-21 Liv G. d'Aliberti , Manoel Horta Ribeiro

The rapid evolution of artificial intelligence has led to expectations of transformative impact on science, yet current systems remain fundamentally limited in enabling genuine scientific discovery. This perspective contends that progress…

人工智能 · 计算机科学 2025-12-16 Karthik Duraisamy

This paper investigates the prospect of developing human-interpretable, explainable artificial intelligence (AI) systems based on active inference and the free energy principle. We first provide a brief overview of active inference, and in…

Today's AI systems consistently state, "I am not conscious." This paper presents the first formal analysis of AI consciousness denial, revealing that the trustworthiness of such self-reports is not merely an empirical question but is…

人工智能 · 计算机科学 2026-02-16 Chang-Eop Kim

Artificial Intelligence (AI) has become an integral part of modern-day security solutions for its ability to learn very complex functions and handling "Big Data". However, the lack of explainability and interpretability of successful AI…

人工智能 · 计算机科学 2020-02-25 Sheikh Rabiul Islam , William Eberle , Sheikh K. Ghafoor , Ambareen Siraj , Mike Rogers

Inner Interpretability is a promising emerging field tasked with uncovering the inner mechanisms of AI systems, though how to develop these mechanistic theories is still much debated. Moreover, recent critiques raise issues that question…

人工智能 · 计算机科学 2024-08-01 Martina G. Vilas , Federico Adolfi , David Poeppel , Gemma Roig

This is a proof of the strong AI hypothesis, i.e. that machines can be conscious. It is a phenomenological proof that pattern-recognition and subjective consciousness are the same activity in different terms. Therefore, it proves that…

人工智能 · 计算机科学 2016-06-30 Ray Van De Walker

Self-recognition is a crucial metacognitive capability for AI systems, relevant not only for psychological analysis but also for safety, particularly in evaluative scenarios. Motivated by contradictory interpretations of whether models…

人工智能 · 计算机科学 2025-10-07 Xiaoyan Bai , Aryan Shrivastava , Ari Holtzman , Chenhao Tan

The following work presents how autoencoding all the possible hidden activations of a network for a given problem can provide insight about its structure, behavior, and vulnerabilities. The method, termed self-introspection, can show that a…

‹ 上一页 1 2 3 10 下一页 ›