English

Performance of Large Language Models in Answering Critical Care Medicine Questions

Computation and Language 2025-09-25 v1

Abstract

Large Language Models have been tested on medical student-level questions, but their performance in specialized fields like Critical Care Medicine (CCM) is less explored. This study evaluated Meta-Llama 3.1 models (8B and 70B parameters) on 871 CCM questions. Llama3.1:70B outperformed 8B by 30%, with 60% average accuracy. Performance varied across domains, highest in Research (68.4%) and lowest in Renal (47.9%), highlighting the need for broader future work to improve models across various subspecialty domains.

Keywords

Cite

@article{arxiv.2509.19344,
  title  = {Performance of Large Language Models in Answering Critical Care Medicine Questions},
  author = {Mahmoud Alwakeel and Aditya Nagori and An-Kwok Ian Wong and Neal Chaisson and Vijay Krishnamoorthy and Rishikesan Kamaleswaran},
  journal= {arXiv preprint arXiv:2509.19344},
  year   = {2025}
}