Question: What is the performance and reasoning ability of OpenAI o1 compared to other large language models in addressing ophthalmology-specific questions? Findings: This study evaluated OpenAI o1 and five LLMs using 6,990 ophthalmological questions from MedMCQA. O1 achieved the highest accuracy (0.88) and macro-F1 score but ranked third in reasoning capabilities based on text-generation metrics. Across subtopics, o1 ranked first in ``Lens'' and ``Glaucoma'' but second to GPT-4o in ``Corneal and External Diseases'', ``Vitreous and Retina'' and ``Oculoplastic and Orbital Diseases''. Subgroup analyses showed o1 performed better on queries with longer ground truth explanations. Meaning: O1's reasoning enhancements may not fully extend to ophthalmology, underscoring the need for domain-specific refinements to optimize performance in specialized fields like ophthalmology.
@article{arxiv.2501.13949,
title = {Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study},
author = {Sahana Srinivasan and Xuguang Ai and Minjie Zou and Ke Zou and Hyunjae Kim and Thaddaeus Wai Soon Lo and Krithi Pushpanathan and Yiming Kong and Anran Li and Maxwell Singer and Kai Jin and Fares Antaki and David Ziyou Chen and Dianbo Liu and Ron A. Adelman and Qingyu Chen and Yih Chung Tham},
journal= {arXiv preprint arXiv:2501.13949},
year = {2025}
}