MIRAGE:基于大型多模态模型的印度普通处方注释的多模态识别与识别
计算机视觉与模式识别
2024-11-13 v2 人工智能
摘要
印度医院仍依赖手写医疗记录,尽管电子医疗记录(EMR)已可用,这使统计分析和记录检索变得复杂。手写记录具有独特挑战,需为训练识别药物及其推荐模式的模型提供专门数据。虽然传统手写体识别方法采用2-D LSTM,但近期研究探索使用多模态大语言模型(MLLM)进行OCR任务。基于此方法,我们专注于从模拟医疗记录中提取药物名称和剂量。我们的 methodology MIRAGE(Multimodal Identification and Recognition of Annotations in indian GEneral prescriptions)涉及在743,118张高分辨率模拟医疗记录图像上对QWEN VL、LLaVA 1.6和Idefics2模型进行微调,这些图像完全来自1,133名印度医生标注。我们的方法在提取药物名称和剂量方面 achieves 82%的准确率。
引用
@article{arxiv.2410.09729,
title = {MIRAGE: Multimodal Identification and Recognition of Annotations in Indian General Prescriptions},
author = {Tavish Mankash and V. S. Chaithanya Kota and Anish De and Praveen Prakash and Kshitij Jadhav},
journal= {arXiv preprint arXiv:2410.09729},
year = {2024}
}
备注
5 pages, 9 figures, 3 tables, submitted to ISBI 2025