English

MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing

Computer Vision and Pattern Recognition 2026-02-09 v1

Abstract

Medical document OCR is challenging due to complex layouts, domain-specific terminology, and noisy annotations, while requiring strict field-level exact matching. Existing OCR systems and general-purpose vision-language models often fail to reliably parse such documents. We propose MeDocVL, a post-trained vision-language model for query-driven medical document parsing. Our framework combines Training-driven Label Refinement to construct high-quality supervision from noisy annotations, with a Noise-aware Hybrid Post-training strategy that integrates reinforcement learning and supervised fine-tuning to achieve robust and precise extraction. Experiments on medical invoice benchmarks show that MeDocVL consistently outperforms conventional OCR systems and strong VLM baselines, achieving state-of-the-art performance under noisy supervision.

Keywords

Cite

@article{arxiv.2602.06402,
  title  = {MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing},
  author = {Wenjie Wang and Wei Wu and Ying Liu and Yuan Zhao and Xiaole Lv and Liang Diao and Zengjian Fan and Wenfeng Xie and Ziling Lin and De Shi and Lin Huang and Kaihe Xu and Hong Li},
  journal= {arXiv preprint arXiv:2602.06402},
  year   = {2026}
}

Comments

20 pages, 8 figures. Technical report

R2 v1 2026-07-01T10:23:45.183Z