English

Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations

Computation and Language 2025-11-14 v2 Artificial Intelligence

Abstract

The rapid growth of medical knowledge and increasing complexity of clinical practice pose challenges. In this context, large language models (LLMs) have demonstrated value; however, inherent limitations remain. Retrieval-augmented generation (RAG) technologies show potential to enhance their clinical applicability. This study reviewed RAG applications in medicine. We found that research primarily relied on publicly available data, with limited application in private data. For retrieval, approaches commonly relied on English-centric embedding models, while LLMs were mostly generic, with limited use of medical-specific LLMs. For evaluation, automated metrics evaluated generation quality and task performance, whereas human evaluation focused on accuracy, completeness, relevance, and fluency, with insufficient attention to bias and safety. RAG applications were concentrated on question answering, report generation, text summarization, and information extraction. Overall, medical RAG remains at an early stage, requiring advances in clinical validation, cross-linguistic adaptation, and support for low-resource settings to enable trustworthy and responsible global use.

Keywords

Cite

@article{arxiv.2511.05901,
  title  = {Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations},
  author = {Rui Yang and Matthew Yu Heng Wong and Huitao Li and Xin Li and Wentao Zhu and Jingchi Liao and Kunyu Yu and Jonathan Chong Kai Liew and Weihao Xuan and Yingjian Chen and Yuhe Ke and Jasmine Chiat Ling Ong and Douglas Teodoro and Chuan Hong and Daniel Shi Wei Ting and Nan Liu},
  journal= {arXiv preprint arXiv:2511.05901},
  year   = {2025}
}