English
Related papers

Related papers: Generalist Foundation Models from a Multimodal Dat…

200 papers

This paper is responding to the MIA-COV19 challenge to classify COVID from non-COVID based on CT lung images. The COVID-19 virus has devastated the world in the last eighteen months by infecting more than 182 million people and causing over…

Image and Video Processing · Electrical Eng. & Systems 2021-07-06 Xiaohong Gao , Yu Qian , Alice Gao

Deep learning methods have demonstrated promising results in predicting BI-RADS scores from mammography images. However, the interpretation of these images can vary, leading to discrepancies even among radiologists. Given the inherent…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Halil Ibrahim Gulluk , Olivier Gevaert

Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current vision-language models are primarily trained on datasets of paired images and reports, not…

Multimodal learning has shown promise in medical imaging, combining complementary modalities like images and text. Vision-language models (VLMs) capture rich diagnostic cues but often require large paired datasets and prompt- or text-based…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Banafsheh Karimian , Giulia Avanzato , Soufian Belharbi , Alexis Guichemerre , Luke McCaffrey , Mohammadhadi Shateri , Eric Granger

Face recognition is a core task in computer vision designed to identify and authenticate individuals by analyzing facial patterns and features. This field intersects with artificial intelligence image processing and machine learning with…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Nhan T. Luu

Medical Visual Question Answering (Med-VQA) holds significant potential for clinical decision support, yet existing efforts primarily focus on 2D imaging with limited task diversity. This paper presents 3D-RAD, a large-scale dataset…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Xiaotang Gai , Jiaxiang Liu , Yichen Li , Zijie Meng , Jian Wu , Zuozhu Liu

Recent advances in clinical AI have enabled remarkable progress across many clinical domains. However, existing benchmarks and models are primarily limited to a small set of modalities and tasks, which hinders the development of large-scale…

Machine Learning · Computer Science 2025-03-21 Wei Dai , Peilin Chen , Malinda Lu , Daniel Li , Haowen Wei , Hejie Cui , Paul Pu Liang

Foundation models have recently gained tremendous popularity in medical image analysis. State-of-the-art methods leverage either paired image-text data via vision-language pre-training or unpaired image data via self-supervised pre-training…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Lei Zhu , Jun Zhou , Rick Siow Mong Goh , Yong Liu

Vision-language models (VLM) have markedly advanced AI-driven interpretation and reporting of complex medical imaging, such as computed tomography (CT). Yet, existing methods largely relegate clinicians to passive observers of final…

Foundation models leveraging vision-language pretraining have shown promise in chest X-ray (CXR) interpretation, yet their real-world performance across diverse populations and diagnostic tasks remains insufficiently evaluated. This study…

Rationale and Objectives: To develop and validate PARROT (Polyglottal Annotated Radiology Reports for Open Testing), a large, multicentric, open-access dataset of fictional radiology reports spanning multiple languages for testing natural…

Computation and Language · Computer Science 2025-08-26 Bastien Le Guellec , Kokou Adambounou , Lisa C Adams , Thibault Agripnidis , Sung Soo Ahn , Radhia Ait Chalal , Tugba Akinci D Antonoli , Philippe Amouyel , Henrik Andersson , Raphael Bentegeac , Claudio Benzoni , Antonino Andrea Blandino , Felix Busch , Elif Can , Riccardo Cau , Armando Ugo Cavallo , Christelle Chavihot , Erwin Chiquete , Renato Cuocolo , Eugen Divjak , Gordana Ivanac , Barbara Dziadkowiec Macek , Armel Elogne , Salvatore Claudio Fanni , Carlos Ferrarotti , Claudia Fossataro , Federica Fossataro , Katarzyna Fulek , Michal Fulek , Pawel Gac , Martyna Gachowska , Ignacio Garcia Juarez , Marco Gatti , Natalia Gorelik , Alexia Maria Goulianou , Aghiles Hamroun , Nicolas Herinirina , Krzysztof Kraik , Dominik Krupka , Quentin Holay , Felipe Kitamura , Michail E Klontzas , Anna Kompanowska , Rafal Kompanowski , Alexandre Lefevre , Tristan Lemke , Maximilian Lindholz , Lukas Muller , Piotr Macek , Marcus Makowski , Luigi Mannacio , Aymen Meddeb , Antonio Natale , Beatrice Nguema Edzang , Adriana Ojeda , Yae Won Park , Federica Piccione , Andrea Ponsiglione , Malgorzata Poreba , Rafal Poreba , Philipp Prucker , Jean Pierre Pruvo , Rosa Alba Pugliesi , Feno Hasina Rabemanorintsoa , Vasileios Rafailidis , Katarzyna Resler , Jan Rotkegel , Luca Saba , Ezann Siebert , Arnaldo Stanzione , Ali Fuat Tekin , Liz Toapanta Yanchapaxi , Matthaios Triantafyllou , Ekaterini Tsaoulia , Evangelia Vassalou , Federica Vernuccio , Johan Wasselius , Weilang Wang , Szymon Urban , Adrian Wlodarczak , Szymon Wlodarczak , Andrzej Wysocki , Lina Xu , Tomasz Zatonski , Shuhang Zhang , Sebastian Ziegelmayer , Gregory Kuchcinski , Keno K Bressem

3D Shape represented as point cloud has achieve advancements in multimodal pre-training to align image and language descriptions, which is curial to object identification, classification, and retrieval. However, the discrete representations…

Computer Vision and Pattern Recognition · Computer Science 2024-02-14 Haoyuan Li , Yanpeng Zhou , Yihan Zeng , Hang Xu , Xiaodan Liang

Training models to apply common-sense linguistic knowledge and visual concepts from 2D images to 3D scene understanding is a promising direction that researchers have only recently started to explore. However, it still remains understudied…

Computer Vision and Pattern Recognition · Computer Science 2023-06-12 Alexandros Delitzas , Maria Parelli , Nikolas Hars , Georgios Vlassis , Sotirios Anagnostidis , Gregor Bachmann , Thomas Hofmann

Automatic radiology reporting has great clinical potential to relieve radiologists from heavy workloads and improve diagnosis interpretation. Recently, researchers have enhanced data-driven neural networks with medical knowledge graphs to…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Mingjie Li , Bingqian Lin , Zicong Chen , Haokun Lin , Xiaodan Liang , Xiaojun Chang

Foundation models trained on large-scale dataset gain a recent surge in CV and NLP. In contrast, development in biomedical domain lags far behind due to data scarcity. To address this issue, we build and release PMC-OA, a biomedical dataset…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Weixiong Lin , Ziheng Zhao , Xiaoman Zhang , Chaoyi Wu , Ya Zhang , Yanfeng Wang , Weidi Xie

Contrastive Language-Image Pre-training (CLIP) is an approach that has advanced research and applications in computer vision, fueling modern recognition systems and generative models. We believe that the main ingredient to the success of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Hu Xu , Saining Xie , Xiaoqing Ellen Tan , Po-Yao Huang , Russell Howes , Vasu Sharma , Shang-Wen Li , Gargi Ghosh , Luke Zettlemoyer , Christoph Feichtenhofer

Detecting misleading patterns in automated diagnostic assistance systems, such as those powered by Artificial Intelligence, is critical to ensuring their reliability, particularly in healthcare. Current techniques for evaluating deep…

Developing a generalist radiology diagnosis system can greatly enhance clinical diagnostics. In this paper, we introduce RadDiag, a foundational model supporting 2D and 3D inputs across various modalities and anatomies, using a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Qiaoyu Zheng , Weike Zhao , Chaoyi Wu , Xiaoman Zhang , Lisong Dai , Hengyu Guan , Yuehua Li , Ya Zhang , Yanfeng Wang , Weidi Xie

Radiology reports are an instrumental part of modern medicine, informing key clinical decisions such as diagnosis and treatment. The worldwide shortage of radiologists, however, restricts access to expert care and imposes heavy workloads,…

The large language model called ChatGPT has drawn extensively attention because of its human-like expression and reasoning abilities. In this study, we investigate the feasibility of using ChatGPT in experiments on using ChatGPT to…

Computation and Language · Computer Science 2024-10-22 Qing Lyu , Josh Tan , Michael E. Zapadka , Janardhana Ponnatapura , Chuang Niu , Kyle J. Myers , Ge Wang , Christopher T. Whitlow
‹ Prev 1 3 4 5 6 7 10 Next ›