中文

indic-punct:面向印度语言的自动标点恢复与逆文本规范化框架

计算与语言 2022-04-01 v1

摘要

自动语音识别(ASR)生成的文本在大多数情况下没有任何标点。文本中标点的缺失会影响可读性。此外,下游 NLP 任务(如情感分析、机器翻译)会因具有标点和句子边界信息而大为受益。我们提出一种利用预训练的 IndicBERT 模型进行文本自动标点的方法。逆文本规范化通过手写加权有限状态转导器(WFST)文法完成。我们已为 11 种印度语言开发了该工具,即 Hindi、Tamil、Telugu、Kannada、Gujarati、Marathi、Odia、Bengali、Assamese、Malayalam 和 Punjabi。所有代码与数据均已公开可用。

关键词

引用

@article{arxiv.2203.16825,
  title  = {indic-punct: An automatic punctuation restoration and inverse text normalization framework for Indic languages},
  author = {Anirudh Gupta and Neeraj Chhimwal and Ankur Dhuriya and Rishabh Gaur and Priyanshi Shah and Harveen Singh Chadha and Vivek Raghavan},
  journal= {arXiv preprint arXiv:2203.16825},
  year   = {2022}
}

备注

Submitted to InterSpeech 2022. arXiv admin note: text overlap with arXiv:2104.05055 by other authors