中文

大规模大语言模型红队测试:应对数学任务中的幻觉问题

计算与语言 2024-01-02 v1 人工智能

摘要

我们考虑了对大语言模型(LLMs)进行基础计算和代数任务红队测试的问题,以评估各种提示技术如何影响输出质量。我们提出了一个框架,用于程序化生成数值问题和谜题,并比较了应用多种红队测试技术与未应用技术时的结果。我们的发现表明,尽管结构化推理和提供逐步示例减缓了答案质量的恶化,但 gpt-3.5-turbo 和 gpt-4 模型并不适合基础计算和推理任务,即使在进行红队测试时也是如此。

关键词

引用

@article{arxiv.2401.00290,
  title  = {Red Teaming for Large Language Models At Scale: Tackling Hallucinations on Mathematics Tasks},
  author = {Aleksander Buszydlik and Karol Dobiczek and Michał Teodor Okoń and Konrad Skublicki and Philip Lippmann and Jie Yang},
  journal= {arXiv preprint arXiv:2401.00290},
  year   = {2024}
}

备注

Accepted to The ART of Safety: Workshop on Adversarial testing and Red-Teaming for generative AI (IJCNLP-AACL 2023)