OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis
Abstract
Dental image analysis plays a pivotal role in supporting accurate diagnosis and treatment planning in oral healthcare. Although recent advances have produced dental AI models for specific tasks and individual imaging modalities, their isolated designs limit practical use in real-world clinical workflows. In this paper, we present OralAgent, the first dental-specialized AI agent that unifies multimodal reasoning, tool-based decision-making, and knowledge-grounded retrieval within an end-to-end automated framework. It integrates 22 visual analysis tools and 368 widely-used classical dental textbooks, enabling autonomous reasoning, planning, tool use, knowledge retrieval, and multi-step workflow execution. Furthermore, we introduce OralCorpus, a large-scale, high-quality bilingual textual resource containing 134.8M tokens curated for dental retrieval-augmented generation (RAG). To evaluate models' multidisciplinary dental knowledge, we construct OralQA-ZH, a Chinese multiple-choice question benchmark consisting of 798 items across eleven oral subspecialties. Extensive experiments demonstrate that OralAgent achieves state-of-the-art performance on the MMOral-Uni, MMOral-OPG, and OralQA-ZH benchmarks, highlighting its effectiveness, interpretability, and adaptability in real-world clinical settings. The code and models are publicly available at https://github.com/isjinghao/OralAgent.
Cite
@article{arxiv.2605.27378,
title = {OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis},
author = {Jing Hao and Siyuan Dai and Yongxin Zhang and Yuci Liang and Jiamin Wu and Jiahao Bao and Yuxuan Fan and Zanting Ye and Yanpeng Sun and Xinyu Zhang and Ming Hu and Liang Zhan and James Kit Hon Tsoi and Linlin Shen and Junjun He and Kuo Feng Hung},
journal= {arXiv preprint arXiv:2605.27378},
year = {2026}
}
Comments
14 pages, 7 figures, 6 tables