衡量法语嵌入对作者风格的敏感性:将文学文本与语言模型改写进行比较
计算与语言
2026-05-12 v1 人工智能
摘要
大型语言模型(LLM)能够令人信服地模仿人类书写风格,但仍不清楚任何语言模型嵌入中编码了多少风格信息以及LLM改写后是否保留了这些信息。我们在法语中进行研究,运用受控的文学数据集量化通过嵌入离散度变化所引起的风格变异效应。我们观察到,嵌入可靠地捕获作者的风格特征,并且这些信号在改写后仍然存在,同时也呈现出LLM特定的模式。这些分析结果为在语言模型时代进行作者模仿检测提供了有前景的方向。
引用
@article{arxiv.2605.10606,
title = {Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings},
author = {Benjamin Icard and Lila Sainero and Alice Breton and Evangelia Zve and Jean-Gabriel Ganascia},
journal= {arXiv preprint arXiv:2605.10606},
year = {2026}
}
备注
To appear in the Proceedings of the 6th International Conference on Natural Language Processing for the Digital Humanities (NLP4DH 2026)