自洫性提升数学推理任务的校准效果
计算与语言
2024-03-18 v1 人工智能
摘要
校准通过建立准确率与模型置信度之间的关联来实现,这对于大型语言模型(LLM)开发至关重要。我们基于自洫性(Wang 等, 2022)设计了三个即插即用校准方法,用于数学推理任务。我们在两个流行基准测试(GSM8K 和 MathQA)上进行评估,使用强大的开源 LLM(Mistral 和 LLaMA2),我们的 methods 较基于 (True) (Kadavath 等, 2022) 或 logits (Kadavath 等, 2022) 的现有方法更好地桥接了模型置信度与准确率。
引用
@article{arxiv.2403.09849,
title = {Self-Consistency Boosts Calibration for Math Reasoning},
author = {Ante Wang and Linfeng Song and Ye Tian and Baolin Peng and Lifeng Jin and Haitao Mi and Jinsong Su and Dong Yu},
journal= {arXiv preprint arXiv:2403.09849},
year = {2024}
}