中文

面向步骤推理的 LLM 潜在空间中的因果概念图

机器学习 2026-04-27 v2 人工智能 统计方法学

摘要

稀疏自编码器可以定位概念在语言模型中的位置,但无法描述在多步骤推理期间它们如何相互作用。我们提出了因果概念图 (Causal Concept Graph, CCG):一种稀疏、可解释的潜在特征上的有向无环图,其中边捕捉学习到的概念之间的因果依赖。我们将任务条件稀疏自编码器用于概念发现,结合 DAGMA 风格的可微结构学习用于图的恢复,并引入因果保真度分数 (Causal Fidelity Score, CFS) 以评估图导向干预是否产生更大的下游效应。实验在 ARC-Challenge、StrategyQA 和 LogiQA 上使用 GPT-2 Medium,经过 15 对 (n=15) 随机种子测试,CCG 的 CFS 为 5.654±0.6255.654\pm0.625,优于 ROME 风格追踪 (3.382±0.2333.382\pm0.233)、仅基于 SAE 的排序 (2.479±0.1962.479\pm0.196) 以及随机基线 (1.032±0.0341.032\pm0.034),在 Bonferroni 校正后 p<0.0001。学习得到的图稀疏(边密度为 5-6%),具有域特定性,且在不同随机种子间稳定。

关键词

引用

@article{arxiv.2603.10377,
  title  = {Causal Concept Graphs in LLM Latent Space for Stepwise Reasoning},
  author = {Md Muntaqim Meherab and Noor Islam S. Mohammad and Faiza Feroz},
  journal= {arXiv preprint arXiv:2603.10377},
  year   = {2026}
}

备注

We have recently encountered author conflicts related to this work and therefore respectfully request the withdrawal of this paper. We believe this step is necessary to address the situation appropriately and maintain academic integrity in the submission