面向混合代码印度社交媒体文本的基于 CRF 的词性标注器
计算与语言
2016-12-26 v1
摘要
在这项工作中,我们描述了一个基于条件随机场(CRF)的系统,用于对混合代码的印度社交媒体文本进行词性(POS)标注,这是我们参加与 2016 年印度理工学院 (BHU) 自然语言处理国际会议同期举行的混合代码印度社交媒体文本词性标注工具竞赛的一部分。我们仅参加了所有三个语言对(孟加拉语-英语、印地语-英语和泰卢固语-英语)的受限模式竞赛。我们的系统实现了 79.99 的总体平均 F1 分数,这是所有参加受限模式竞赛的 16 个系统中最高的总体平均 F1 分数。
引用
@article{arxiv.1612.07956,
title = {A CRF Based POS Tagger for Code-mixed Indian Social Media Text},
author = {Kamal Sarkar},
journal= {arXiv preprint arXiv:1612.07956},
year = {2016}
}
备注
This work is awarded the first prize in the NLP tool contest on "POS Tagging for Code-Mixed Indian Social Media Text", held in conjunction with the 13th International Conference on Natural Language Processing 2016(ICON 2016), Indian Institute of Technology (BHU), India