用于检测推文中药物提及的深度神经网络集成方法
摘要
目的:经过多年研究,Twitter 帖子现已被视为患者生成数据的重要来源,为人群健康提供独特洞见。将 Twitter 数据纳入药物流行病学研究的基本步骤是自动识别推文中的药物提及。鉴于基于词典搜索药物名称可能因拼写错误或与常见词歧义而失效,我们提出一种更先进的方法进行识别。方法:我们提出 Kusuri,一种集成学习分类器,能够识别提及药品与膳食补充剂的推文。Kusuri(日语中“药物”之意)由两个模块组成。首先,四种不同分类器(基于词典、基于拼写变体、基于模式以及一种基于弱训练神经网络)并行应用以发现可能含药物名的推文。其次,使用编码所发现推文中重要词的形态、语义与长距离依赖的深度神经网络集成来做出最终判定。结果:在一个平衡的(50-50)含 15,005 条推文的语料上,Kusuri 表现出接近人工标注者的性能,F1 值达 93.7%,系该语料迄今最佳分数。在一个由 113 名 Twitter 用户发布的所有推文构成的语料(98,959 条推文,仅 0.26% 提及药物)上,Kusuri 获得 76.3% 的 F1 值。此前并无在如此极度不平衡数据集上运行的同类药物抽取系统。结论:该系统识别提及药物名推文的性能足以保障其实用性,并可供集成至更大型自然语言处理系统中。
引用
@article{arxiv.1904.05308,
title = {Deep Neural Networks Ensemble for Detecting Medication Mentions in Tweets},
author = {Davy Weissenbacher and Abeed Sarker and Ari Klein and Karen O'Connor and Arjun Magge Ranganatha and Graciela Gonzalez-Hernandez},
journal= {arXiv preprint arXiv:1904.05308},
year = {2019}
}
备注
This is a pre-copy-editing, author-produced PDF of an article accepted for publication in JAMIA following peer review. The definitive publisher-authenticated version is "D. Weissenbacher, A. Sarker, A. Klein, K. O'Connor, A. Magge, G. Gonzalez-Hernandez, Deep neural networks ensemble for detecting medication mentions in tweets, Journal of the American Medical Informatics Association, ocz156, 2019"