English

Performance Evaluation of Open-Source Large Language Models for Assisting Pathology Report Writing in Japanese

Computation and Language 2026-03-13 v1 Artificial Intelligence

Abstract

The performance of large language models (LLMs) for supporting pathology report writing in Japanese remains unexplored. We evaluated seven open-source LLMs from three perspectives: (A) generation and information extraction of pathology diagnosis text following predefined formats, (B) correction of typographical errors in Japanese pathology reports, and (C) subjective evaluation of model-generated explanatory text by pathologists and clinicians. Thinking models and medical-specialized models showed advantages in structured reporting tasks that required reasoning and in typo correction. In contrast, preferences for explanatory outputs varied substantially across raters. Although the utility of LLMs differed by task, our findings suggest that open-source LLMs can be useful for assisting Japanese pathology report writing in limited but clinically relevant scenarios.

Keywords

Cite

@article{arxiv.2603.11597,
  title  = {Performance Evaluation of Open-Source Large Language Models for Assisting Pathology Report Writing in Japanese},
  author = {Masataka Kawai and Singo Sakashita and Shumpei Ishikawa and Shogo Watanabe and Anna Matsuoka and Mikio Sakurai and Yasuto Fujimoto and Yoshiyuki Takahara and Atsushi Ohara and Hirohiko Miyake and Genichiro Ishii},
  journal= {arXiv preprint arXiv:2603.11597},
  year   = {2026}
}

Comments

9 pages (including bibliography), 2 figures, 6 tables