PARROT: An Open Multilingual Radiology Reports Dataset
Abstract
Rationale and Objectives: To develop and validate PARROT (Polyglottal Annotated Radiology Reports for Open Testing), a large, multicentric, open-access dataset of fictional radiology reports spanning multiple languages for testing natural language processing applications in radiology. Materials and Methods: From May to September 2024, radiologists were invited to contribute fictional radiology reports following their standard reporting practices. Contributors provided at least 20 reports with associated metadata including anatomical region, imaging modality, clinical context, and for non-English reports, English translations. All reports were assigned ICD-10 codes. A human vs. AI report differentiation study was conducted with 154 participants (radiologists, healthcare professionals, and non-healthcare professionals) assessing whether reports were human-authored or AI-generated. Results: The dataset comprises 2,658 radiology reports from 76 authors across 21 countries and 13 languages. Reports cover multiple imaging modalities (CT: 36.1%, MRI: 22.8%, radiography: 19.0%, ultrasound: 16.8%) and anatomical regions, with chest (19.9%), abdomen (18.6%), head (17.3%), and pelvis (14.1%) being most prevalent. In the differentiation study, participants achieved 53.9% accuracy (95% CI: 50.7%-57.1%) in distinguishing between human and AI-generated reports, with radiologists performing significantly better (56.9%, 95% CI: 53.3%-60.6%, p<0.05) than other groups. Conclusion: PARROT represents the largest open multilingual radiology report dataset, enabling development and validation of natural language processing applications across linguistic, geographic, and clinical boundaries without privacy constraints.
Cite
@article{arxiv.2507.22939,
title = {PARROT: An Open Multilingual Radiology Reports Dataset},
author = {Bastien Le Guellec and Kokou Adambounou and Lisa C Adams and Thibault Agripnidis and Sung Soo Ahn and Radhia Ait Chalal and Tugba Akinci D Antonoli and Philippe Amouyel and Henrik Andersson and Raphael Bentegeac and Claudio Benzoni and Antonino Andrea Blandino and Felix Busch and Elif Can and Riccardo Cau and Armando Ugo Cavallo and Christelle Chavihot and Erwin Chiquete and Renato Cuocolo and Eugen Divjak and Gordana Ivanac and Barbara Dziadkowiec Macek and Armel Elogne and Salvatore Claudio Fanni and Carlos Ferrarotti and Claudia Fossataro and Federica Fossataro and Katarzyna Fulek and Michal Fulek and Pawel Gac and Martyna Gachowska and Ignacio Garcia Juarez and Marco Gatti and Natalia Gorelik and Alexia Maria Goulianou and Aghiles Hamroun and Nicolas Herinirina and Krzysztof Kraik and Dominik Krupka and Quentin Holay and Felipe Kitamura and Michail E Klontzas and Anna Kompanowska and Rafal Kompanowski and Alexandre Lefevre and Tristan Lemke and Maximilian Lindholz and Lukas Muller and Piotr Macek and Marcus Makowski and Luigi Mannacio and Aymen Meddeb and Antonio Natale and Beatrice Nguema Edzang and Adriana Ojeda and Yae Won Park and Federica Piccione and Andrea Ponsiglione and Malgorzata Poreba and Rafal Poreba and Philipp Prucker and Jean Pierre Pruvo and Rosa Alba Pugliesi and Feno Hasina Rabemanorintsoa and Vasileios Rafailidis and Katarzyna Resler and Jan Rotkegel and Luca Saba and Ezann Siebert and Arnaldo Stanzione and Ali Fuat Tekin and Liz Toapanta Yanchapaxi and Matthaios Triantafyllou and Ekaterini Tsaoulia and Evangelia Vassalou and Federica Vernuccio and Johan Wasselius and Weilang Wang and Szymon Urban and Adrian Wlodarczak and Szymon Wlodarczak and Andrzej Wysocki and Lina Xu and Tomasz Zatonski and Shuhang Zhang and Sebastian Ziegelmayer and Gregory Kuchcinski and Keno K Bressem},
journal= {arXiv preprint arXiv:2507.22939},
year = {2025}
}
Comments
Corrected affiliations (no change to the paper)