CLEAR:对比专家与业余模型的文本反馈以提升推理能力
计算与语言
2025-04-11 v1
摘要
我们介绍 CLEAR(Contrasting Textual Feedback with Experts and Amateurs for Reasoning),一种新颖的语言模型推理方法,利用较大(专家)模型和较小(业余)模型的优势。专家模型和业余模型各自对模型初始输出提供反馈,并相互对比以生成细化后的反馈。随后将此反馈应用于迭代式改进 CLEAR 的响应。我们的实验表明,CLEAR 在多个具有挑战性的推理任务中超越了最先进的方法,包括故事提纲改进(有趣度提升最高达 19.6%)、受约束生成(覆盖率提升最高达 18.5%)、数学推理(准确率提升最高达 6.7%)以及降低有害输出(有害程度降低最高达 22%)。
引用
@article{arxiv.2504.07116,
title = {CLEAR: Contrasting Textual Feedback with Experts and Amateurs for Reasoning},
author = {Andrew Rufail and Daniel Kim and Sean O'Brien and Kevin Zhu},
journal= {arXiv preprint arXiv:2504.07116},
year = {2025}
}
备注
Accepted at the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), Student Research Workshop (SRW)