中文

巴西法律文档的序列感知多模态页面分类

计算与语言 2022-07-18 v2

摘要

巴西最高法院每学期收到数万起案件。法院员工花费数千小时执行这些案件的初步分析与分类——这使其精力从案件管理工作流中后续更复杂的阶段转移。本文中,我们探索来自巴西最高法院文档的多模态分类。我们在一个包含 6,510 份诉讼(339,478 页)的新型多模态数据集上训练和评估我们的方法,该数据集带有将每页手动标注为六类之一的标注。每份诉讼是一个有序的页面序列,以图像形式存储,并通过光学字符识别提取相应文本。我们首先训练两个单模态分类器:在 ImageNet 上预训练的 ResNet 在图像上微调,以及具有多核大小滤波器的卷积网络从零开始在文档文本上训练。我们将它们用作视觉和文本特征的提取器,然后通过我们提出的融合模块进行组合。我们的融合模块可通过使用缺失数据的学习嵌入来处理缺失的文本或视觉输入。此外,我们尝试使用双向长短期记忆(biLSTM)网络和线性链条件随机场来建模页面的序列性质。多模态方法优于文本和视觉分类器,尤其是在利用页面序列性质时。

关键词

引用

@article{arxiv.2207.00748,
  title  = {Sequence-aware multimodal page classification of Brazilian legal documents},
  author = {Pedro H. Luz de Araujo and Ana Paula G. S. de Almeida and Fabricio A. Braz and Nilton C. da Silva and Flavio de Barros Vidal and Teofilo E. de Campos},
  journal= {arXiv preprint arXiv:2207.00748},
  year   = {2022}
}

备注

11 pages, 6 figures. This preprint, which was originally written on 8 April 2021, has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this article is published in the International Journal on Document Analysis and Recognition, and is available online at https://doi.org/10.1007/s10032-022-00406-7 and https://rdcu.be/cRvvV