LexSumm 与 LexT5:面向英语法律摘要任务的基准测试与建模
计算与语言
2024-10-15 v1
摘要
在不断演变的 NLP 领域中,基准测试作为衡量进展的标尺。然而,现有的 Legal NLP 基准测试仅关注预测任务,忽略了生成任务。本工作构建 LexSumm,一个用于评估英语法律摘要任务的基准测试。其包含来自美国、英国、欧盟和印度等不同司法管辖区的八个英语法律摘要数据集。此外,我们发布了 LexT5,一个面向法律领域的序列到序列模型,解决了现有 BERT 系列 encoder-only 模型在法律领域的局限性。我们通过在 LegalLAMA 上的 zero-shot probing 和在 LexSumm 上的 fine-tuning 评估了其能力。我们的分析表明,即使是基于 zero-shot LLM 生成的摘要,也存在抽象性和忠实性错误,表明仍有进一步改进的潜力。LexSumm 基准测试和 LexT5 模型已提供:https://github.com/TUMLegalTech/LexSumm-LexT5。
引用
@article{arxiv.2410.09527,
title = {LexSumm and LexT5: Benchmarking and Modeling Legal Summarization Tasks in English},
author = {T. Y. S. S. Santosh and Cornelius Weiss and Matthias Grabmair},
journal= {arXiv preprint arXiv:2410.09527},
year = {2024}
}
备注
Accepted to NLLP Workshop, EMNLP 2024