大型语言模型的风格计量学应用
计算与语言
2025-10-28 v1 数字图书馆
摘要
我们展示了大型语言模型(LLM)可用于区分不同作者的写作风格.具体而言,训练于单个作者作品的GPT-2模型,会比预测来自其他作者的 held-out文本,更准确地预测该作者的 held-out文本.我们认为,通过这种方式,在单个作者作品上训练的模型体现了该作者独特的写作风格.我们首先在由八位不同(已知)作者所写的书籍上演示了该方法.我们还使用该方法确认了R.P. Thompson关于其署名权的 authorship,后者原本归属于著名的15卷奥斯系列之书,原归F.L.Baum所有.
引用
@article{arxiv.2510.21958,
title = {A Stylometric Application of Large Language Models},
author = {Harrison F. Stropkay and Jiayi Chen and Mohammad J. Latifi and Daniel N. Rockmore and Jeremy R. Manning},
journal= {arXiv preprint arXiv:2510.21958},
year = {2025}
}
备注
All code and data needed to reproduce the results in this paper are available at https://github.com/ContextLab/llm-stylometry