中文

大型语言模型在系统综述中的有效性

计算与语言 2024-10-29 v2 机器学习

摘要

本研究通过对 ESG(环境、社会、治理)因素与财务绩效关系的系统综述,探讨了大型语言模型(LLM)在解读既有文献方面的有效性。主要目标是评估 LLM 如何在一套 ESG 主题论文语料库上复制系统综述。我们编译并人工编码了一个包含 88 篇2020年3月至2024年5月发表的相关论文的数据库。此外,我们使用了一套来自2015年1月至2020年2月的 238 篇 ESG 文献综述论文。我们对 Meta AI 的 Llama 3 8B 和 OpenAI 的 GPT-4o 两种当前主流 LLM,评估其相对于人工分类在两套论文上的解释准确性。随后,我们将这些结果与一个“自定义 GPT”和使用 238 篇论文作为训练数据的微调 GPT-4o Mini 模型进行比较。微调 GPT-4o Mini 模型在 prompt 1 的总体准确率上平均提升了 28.3%。而“自定义 GPT”在 prompt 2 和 prompt 3 的总体准确率上分别平均提升了 3.0% 和 15.7%。我们的发现表明,投资者和机构可利用 LLM 对 ESG 投资相关的复杂证据进行概括,从而实现更快的决策并提高市场效率。

关键词

引用

@article{arxiv.2408.04646,
  title  = {Efficacy of Large Language Models in Systematic Reviews},
  author = {Aaditya Shah and Shridhar Mehendale and Siddha Kanthi},
  journal= {arXiv preprint arXiv:2408.04646},
  year   = {2024}
}

备注

Both Shah and Mehendale contributed equally to this work; order of authorship is random. This paper will be published in the proceedings of The 2nd International Conference on Foundation and Large Language Models (FLLM2024) in IEEE Xplore