中文

大型语言模型的风格计量学应用

计算与语言 2025-10-28 v1 数字图书馆

摘要

我们展示了大型语言模型(LLM)可用于区分不同作者的写作风格.具体而言,训练于单个作者作品的GPT-2模型,会比预测来自其他作者的 held-out文本,更准确地预测该作者的 held-out文本.我们认为,通过这种方式,在单个作者作品上训练的模型体现了该作者独特的写作风格.我们首先在由八位不同(已知)作者所写的书籍上演示了该方法.我们还使用该方法确认了R.P. Thompson关于其署名权的 authorship,后者原本归属于著名的15卷奥斯系列之书,原归F.L.Baum所有.

关键词

引用

@article{arxiv.2510.21958,
  title  = {A Stylometric Application of Large Language Models},
  author = {Harrison F. Stropkay and Jiayi Chen and Mohammad J. Latifi and Daniel N. Rockmore and Jeremy R. Manning},
  journal= {arXiv preprint arXiv:2510.21958},
  year   = {2025}
}

备注

All code and data needed to reproduce the results in this paper are available at https://github.com/ContextLab/llm-stylometry