English

Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models

Computation and Language 2026-05-21 v1 Computers and Society

Abstract

Medical large language models (LLMs), including custom medical GPTs (MedGPTs) and open-source models, are increasingly deployed on web platforms to provide clinical guidance. However, they pose risks of hallucination, policy noncompliance, and unsafe design. We conduct a large-scale assessment of 6,233 MedGPTs, evaluating a stratified sample of 1,500, together with 10 open-source LLMs. We introduce two frameworks: MedGPT-HEval for hallucination detection and an LLM-based pipeline for assessing policy violations and developer intent. Our results show that 25-30% of MedGPTs exhibit low factual accuracy, with bottom- and middle-tier models at highest risk; 33.6-54.3% violate operational thresholds, and 57.06% of Action-enabled models lack adequate privacy disclosures. Compared with open-source models, MedGPTs achieve higher factual accuracy and semantic alignment, though open-source models are more stable. These results reveal systemic gaps in hallucination and compliance, highlighting the need for multi-metric evaluation and stronger safeguards. We release HAA-MedGPT, a structured dataset that supports future research on the safety of web-facing medical LLMs.

Keywords

Cite

@article{arxiv.2605.20591,
  title  = {Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models},
  author = {Sunday Oyinlola Ogundoyin and Muhammad Ikram and Rahat Masood},
  journal= {arXiv preprint arXiv:2605.20591},
  year   = {2026}
}
R2 v1 2026-07-22T07:23:00.501Z