链路思考推理的神经语言模型表征能力
计算与语言
2025-01-27 v2 形式语言与自动机理论
摘要
现代语言模型 (LMs) 的性能通过链路思考 (CoT) 推理得到提升,即生成引导模型走向最终答案的中间结果。这种提升的一个可能解释是,CoT 推理扩展了 LM 的计算能力,已知具有额外存储空间的 RNN 和 Transformer 是图灵完备的。将 LM 与图灵机进行比较,然而会引入范畴错误——图灵机决定语言成员资格,而 LM 定义字符串上的分布。为弥合这一差距,我们在概率语境中形式化 CoT 推理。我们提出了若干结果,表明具有 CoT 推理的循环和 Transformer LM 能够表示与概率图灵机相同的 string 分布家族。
引用
@article{arxiv.2406.14197,
title = {On the Representational Capacity of Neural Language Models with Chain-of-Thought Reasoning},
author = {Franz Nowak and Anej Svete and Alexandra Butoi and Ryan Cotterell},
journal= {arXiv preprint arXiv:2406.14197},
year = {2025}
}
备注
Published at ACL 2024