不寻常的对决:通过独特情景评估 LLM 的创意写作
计算与语言
2024-06-25 v1
摘要
本文是论文 "A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing" 的总结,后者发表于 EMNLP 2023 的 Findings。我们评估了一系列最新的、经过指令微调的大语言模型(LLMs)在英语创意写作任务上的表现,并将其与人类作家进行比较。为此,我们使用特定量身定制的提示(基于 Ignatius J. Reilly——约翰·肯尼·图尔《卧龙凤雄》主角——与恐龙的史诗对决)来最小化训练数据泄露的风险,并迫使模型进行创意而非重用现有故事。相同的提示被呈现给 LLMs 和人类作家,评估由人类进行,使用详细的评分标准,包括流畅性、风格、原创性或幽默感等各方面。结果显示,一些最新的商业 LLMs 在大多数评估维度上与人类作家相当或略胜一筑。开源 LLMs 落后于人类。人类在原创性方面保持领先,只有前三名 LLMs 能够以接近人类的水平处理幽默。
引用
@article{arxiv.2406.15891,
title = {The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario},
author = {Carlos Gómez-Rodríguez and Paul Williams},
journal= {arXiv preprint arXiv:2406.15891},
year = {2024}
}
备注
Published in the XIX Conference of the Spanish Association for Artificial Intelligence (CAEPIA), 2024. Summary of our paper "A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing", published in Findings of EMNLP