English

Automated HIV Screening on Dutch Electronic Health Records with Large Language Models

Computation and Language 2025-10-28 v2

Abstract

Efficient screening and early diagnosis of HIV are critical for reducing onward transmission. Although large scale laboratory testing is not feasible, the widespread adoption of Electronic Health Records (EHRs) offers new opportunities to address this challenge. Existing research primarily focuses on applying machine learning methods to structured data, such as patient demographics, for improving HIV diagnosis. However, these approaches often overlook unstructured text data such as clinical notes, which potentially contain valuable information relevant to HIV risk. In this study, we propose a novel pipeline that leverages a Large Language Model (LLM) to analyze unstructured EHR text and determine a patient's eligibility for further HIV testing. Experimental results on clinical data from Erasmus University Medical Center Rotterdam demonstrate that our pipeline achieved high accuracy while maintaining a low false negative rate.

Keywords

Cite

@article{arxiv.2510.19879,
  title  = {Automated HIV Screening on Dutch Electronic Health Records with Large Language Models},
  author = {Lang Zhou and Amrish Jhingoer and Yinghao Luo and Klaske Vliegenthart--Jongbloed and Carlijn Jordans and Ben Werkhoven and Tom Seinen and Erik van Mulligen and Casper Rokx and Yunlei Li},
  journal= {arXiv preprint arXiv:2510.19879},
  year   = {2025}
}

Comments

28 pages, 6 figures