中文

利用数据增强与语言模型集成从Twitter抽取药物名称

计算与语言 2021-11-15 v1 机器学习

摘要

BioCreative VII Track 3挑战赛聚焦于在Twitter用户时间线中识别药物名称。对于我们向该挑战赛的提交,我们通过使用多种数据增强技术扩展了可用的训练数据。随后利用增强后的数据对已在通用领域Twitter内容上预训练的语言模型集成进行微调。所提出的方法优于先前的state-of-the-art算法Kusuri,并在我们所选目标函数重叠F1分数(overlapping F1 score)的竞赛中排名靠前。

关键词

引用

@article{arxiv.2111.06664,
  title  = {Extraction of Medication Names from Twitter Using Augmentation and an Ensemble of Language Models},
  author = {Igor Kulev and Berkay Köprü and Raul Rodriguez-Esteban and Diego Saldana and Yi Huang and Alessandro La Torraca and Elif Ozkirimli},
  journal= {arXiv preprint arXiv:2111.06664},
  year   = {2021}
}

备注

Proceedings of the BioCreative VII Challenge Evaluation Workshop