CSSL-MHTR:面向可扩展多文字手写文本识别的持续自监督学习
计算机视觉与模式识别
2024-04-30 v2
摘要
自监督学习近来已成为文档分析中的一个有力替代方案。这些方法现在能够学习高质量的图像表示,并克服需要大量标注数据的监督方法的局限性。然而,这些方法无法以增量方式捕获新知识,即数据按顺序呈现给模型,这更接近真实场景。在本文中,我们探索持续自监督学习的潜力,以缓解手写文本识别(作为序列识别的一个例子)中的灾难性遗忘问题。我们的方法包括为每个任务添加称为适配器的中间层,并在学习当前任务的同时高效地从先前模型中蒸馏知识。我们提出的框架在计算和内存复杂度上均高效。为证明其有效性,我们通过将学习到的模型迁移到多样的文本识别下游任务(包括拉丁和非拉丁文字)来评估我们的方法。据我们所知,这是持续自监督学习在手写文本识别中的首次应用。我们在英、意、俄三种文字上取得了最先进的(SOTA)性能,同时每个任务仅增加少量参数。代码与训练好的模型将公开可用。
引用
@article{arxiv.2303.09347,
title = {CSSL-MHTR: Continual Self-Supervised Learning for Scalable Multi-script Handwritten Text Recognition},
author = {Marwa Dhiaf and Mohamed Ali Souibgui and Kai Wang and Yuyang Liu and Yousri Kessentini and Alicia Fornés and Ahmed Cheikh Rouhou},
journal= {arXiv preprint arXiv:2303.09347},
year = {2024}
}
备注
Due to current company policy constraints, we are compelled to withdraw our paper. The organization's guidelines prohibit us from proceeding with the publication of this work at this time. We apologize for any inconvenience this may cause and appreciate your understanding in this matter