DALPHIN: Benchmarking Digital Pathology AI Copilots Against Pathologists on an Open Multicentric Dataset
Abstract
Foundation models with visual question answering capabilities for digital pathology are emerging. Such unprecedented technology requires independent benchmarking to assess its potential in assisting pathologists in routine diagnostics. We created DALPHIN, the first multicentric open benchmark for pathology AI copilots, comprising 1236 images from 300 cases, spanning 130 rare to common diagnoses, 6 countries, and 14 subspecialties. The DALPHIN design and dataset are introduced alongside a human performance benchmark of 31 pathologists from 10 countries with varying expertise. We report results for two general-purpose (GPT-5, Gemini 2.5 Pro) and one pathology-specific copilot (PathChat+) for sequential and independent answer generation. We observed no statistically significant difference from expert-level performance in four of six tasks for PathChat, 2/6 tasks for Gemini, and 1/6 tasks for GPT. DALPHIN is publicly released with sequestered, indirectly accessible ground truth to foster robust and enduring benchmarking. Data, methods, and the evaluation platform are accessible through dalphin.grand-challenge.org.
Cite
@article{arxiv.2605.03544,
title = {DALPHIN: Benchmarking Digital Pathology AI Copilots Against Pathologists on an Open Multicentric Dataset},
author = {Carlijn Lems and Sander Moonemans and Natálie Klubíčková and Biagio Brattoli and Taebum Lee and Seokhwi Kim and Veronica Vilaplana and Laura Pons and Sapir Hochman and Mauricio Eduardo Suárez-Franck and Pedro Luis Fernandez and Julius Drachneris and Donatas Petroska and Renaldas Augulis and Arvydas Laurinavicius and Domingos Oliveira and Diana Montezuma and Anouk B. Bouwmeester and Dominique van Midden and Anne-Marie Vos and Shoko Vos and Jolique van Ipenburg and Maschenka Balkenhol and Koen Winkler and Iris Nagtegaal and Konnie Hebeda and Uta Flucke and Katrien Grünberg and Josef Skopal and Brinder S. Chohan and Jordi Temprana-Salvador and Enrico Munari and Luca Cima and Giulia Querzoli and Yosamin Gonzalez Belisario and Jaeike W. Faber and Geert J. L. H. van Leenders and Jan H. von der Thüsen and Lodewijk A. A. Brosens and Ronald R. de Krijger and Pieter Wesseling and Sandrine Florquin and Mateusz Maniewski and Adam Kowalewski and Robert Barna and Dina Tiniakos and Joan Lop Gros and Rogier Donders and Jake S. F. Maurits and Ming Yang Lu and Chengkuan Chen and Faisal Mahmood and Jeroen van der Laak and Nadieh Khalili and Frédérique Meeuwsen and Francesco Ciompi},
journal= {arXiv preprint arXiv:2605.03544},
year = {2026}
}
Comments
Our dataset is available at https://zenodo.org/records/18609450 , our code is available at https://github.com/computationalpathologygroup/DALPHIN , and our benchmark is available at https://dalphin.grand-challenge.org/