indic-punct:面向印度语言的自动标点恢复与逆文本规范化框架
计算与语言
2022-04-01 v1
摘要
自动语音识别(ASR)生成的文本在大多数情况下没有任何标点。文本中标点的缺失会影响可读性。此外,下游 NLP 任务(如情感分析、机器翻译)会因具有标点和句子边界信息而大为受益。我们提出一种利用预训练的 IndicBERT 模型进行文本自动标点的方法。逆文本规范化通过手写加权有限状态转导器(WFST)文法完成。我们已为 11 种印度语言开发了该工具,即 Hindi、Tamil、Telugu、Kannada、Gujarati、Marathi、Odia、Bengali、Assamese、Malayalam 和 Punjabi。所有代码与数据均已公开可用。
引用
@article{arxiv.2203.16825,
title = {indic-punct: An automatic punctuation restoration and inverse text normalization framework for Indic languages},
author = {Anirudh Gupta and Neeraj Chhimwal and Ankur Dhuriya and Rishabh Gaur and Priyanshi Shah and Harveen Singh Chadha and Vivek Raghavan},
journal= {arXiv preprint arXiv:2203.16825},
year = {2022}
}
备注
Submitted to InterSpeech 2022. arXiv admin note: text overlap with arXiv:2104.05055 by other authors