利用数据增强与语言模型集成从Twitter抽取药物名称
计算与语言
2021-11-15 v1 机器学习
摘要
BioCreative VII Track 3挑战赛聚焦于在Twitter用户时间线中识别药物名称。对于我们向该挑战赛的提交,我们通过使用多种数据增强技术扩展了可用的训练数据。随后利用增强后的数据对已在通用领域Twitter内容上预训练的语言模型集成进行微调。所提出的方法优于先前的state-of-the-art算法Kusuri,并在我们所选目标函数重叠F1分数(overlapping F1 score)的竞赛中排名靠前。
引用
@article{arxiv.2111.06664,
title = {Extraction of Medication Names from Twitter Using Augmentation and an Ensemble of Language Models},
author = {Igor Kulev and Berkay Köprü and Raul Rodriguez-Esteban and Diego Saldana and Yi Huang and Alessandro La Torraca and Elif Ozkirimli},
journal= {arXiv preprint arXiv:2111.06664},
year = {2021}
}
备注
Proceedings of the BioCreative VII Challenge Evaluation Workshop