中文

机器翻译中未识别词的消解

计算与语言 2010-07-29 v2

摘要

本文提出了一种在基于文本的机器翻译(TBMT)中消解未识别词汇单元的机制。在机器翻译(MT)系统中,拥有完整的词典是不太可能的,因此迫切需要一种新的机制来处理未识别词的问题。这些未知词可能是缩写、名称、首字母缩略词和新引入的术语。我们提出了一种用于消解未识别词的算法。该算法以语篇单元(原始语篇)为分析单元,并提供对词典的实时更新。我们已将算法手动应用于报纸片段。除了回指和预指消解外,许多未知词,特别是名称和缩写,都被更新到了词典中。

关键词

引用

@article{arxiv.0911.1517,
  title  = {Resolution of Unidentified Words in Machine Translation},
  author = {Sana Ullah and M. Asdaque Hussain and Kyung Sup Kwak},
  journal= {arXiv preprint arXiv:0911.1517},
  year   = {2010}
}

备注

4 pages, 2 figures, 2 tables, The 4th annual International New Exploratory Technologies Conference 2007 (NEXT 2007), Seoul, South Korea