大语言模型用于总结捷克语历史文献及其扩展应用
计算与语言
2025-08-15 v1
摘要
文本概述是将较大的文本简化为保留其核心含义和关键信息的简洁版本的任务。虽然概述已在英语及其他高资源语言中得到广泛探索,但捷克语概述,特别是针对历史文献,仍属underexplored 领域,由于语言复杂性和缺乏标注数据集。大型语言模型如 Mistral 和 mT5 在许多自然语言处理任务和语言上表现出色。因此,我们在捷克语概述任务中采用这些模型,取得两项关键成果:(1)在捷克语概述数据集 SumeCzech 上取得新 state-of-the-art 结果,(2)引入 novel 数据集 Posel od Čerchova,用于历史捷克语文献概述,提供 baseline 结果。正是由于这些贡献,为捷克语文本概述的进一步发展开辟了广阔的潜力,并为捷克语历史文本处理的研究指明了新方向。
引用
@article{arxiv.2508.10368,
title = {Large Language Models for Summarizing Czech Historical Documents and Beyond},
author = {Václav Tran and Jakub Šmíd and Jiří Martínek and Ladislav Lenc and Pavel Král},
journal= {arXiv preprint arXiv:2508.10368},
year = {2025}
}
备注
Published in Proceedings of the 17th International Conference on Agents and Artificial Intelligence - Volume 2 (ICAART 2025). Official version: https://www.scitepress.org/Link.aspx?doi=10.5220/0013374100003890