English
Related papers

Related papers: MedGemma 1.5 Technical Report

200 papers

While 3D visual self-supervised learning (vSSL) shows promising results in capturing visual representations, it overlooks the clinical knowledge from radiology reports. Meanwhile, 3D medical vision-language pre-training (MedVLP) remains…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Che Liu , Cheng Ouyang , Yinda Chen , Cesar César Quilodrán-Casas , Lei Ma , Jie Fu , Yike Guo , Anand Shah , Wenjia Bai , Rossella Arcucci

Breast cancer remains a critical global health challenge, necessitating early and accurate detection for effective treatment. This paper introduces a methodology that combines automated image augmentation selection (RandAugment) with search…

Image and Video Processing · Electrical Eng. & Systems 2023-11-21 Leon Hamnett , Mary Adewunmi , Modinat Abayomi , Kayode Raheem , Fahad Ahmed

Medical image segmentation is critical for clinical diagnosis, treatment planning, and monitoring, yet segmentation models often struggle with uncertainties stemming from occlusions, ambiguous boundaries, and variations in imaging devices.…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Xiao Ma , Yuhui Tao , Zetian Zhang , Yuhan Zhang , Xi Wang , Sheng Zhang , Zexuan Ji , Yizhe Zhang , Qiang Chen , Guang Yang

The Medical Segment Anything Model (MedSAM) has shown remarkable performance in medical image segmentation, drawing significant attention in the field. However, its sensitivity to varying prompt types and locations poses challenges. This…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Nan Zhou , Ke Zou , Kai Ren , Mengting Luo , Linchao He , Meng Wang , Yidi Chen , Yi Zhang , Hu Chen , Huazhu Fu

Medical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. MedVLP can provide useful features to downstream tasks and…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Yang Zhou , Tan Li Hui Faith , Yanyu Xu , Sicong Leng , Xinxing Xu , Yong Liu , Rick Siow Mong Goh

Recent advances in artificial intelligence (AI), in particular self-supervised learning of foundation models (FMs), are revolutionizing medical imaging and computational pathology (CPath). A constant challenge in the analysis of digital…

Large Language Models (LLMs) have demonstrated impressive capabilities across natural language processing tasks. However, their application to specialized domains such as medicine and biology requires further optimization to ensure factual…

Computation and Language · Computer Science 2025-02-06 Seonok Kim

While the Segment Anything Model (SAM) excels in semantic segmentation for general-purpose images, its performance significantly deteriorates when applied to medical images, primarily attributable to insufficient representation of medical…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Yiming Zhang , Tianang Leng , Kun Han , Xiaohui Xie

The global demand for radiologists is increasing rapidly due to a growing reliance on medical imaging services, while the supply of radiologists is not keeping pace. Advances in computer vision and image processing technologies present…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Shehroz S. Khan , Petar Przulj , Ahmed Ashraf , Ali Abedi

Deep learning-based medical image segmentation is increasingly used to support clinical diagnosis and develop new treatment strategies. However, model performance remains limited by the scarcity of high-quality annotated data and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Nathan Molinier , Hendrik Möller , Thomas Dagonneau , Anna Curto-Vilalta , Robert Graf , Matan Atad , Daniel Rueckert , Jan S. Kirschke , Julien Cohen-Adad

Omni-tomography is enabled by interior tomography that has been developed over the past five years. By omni-tomography, we envision that the next stage of biomedical imaging will be the grand fusion of many tomographic modalities into a…

Medical Physics · Physics 2012-12-24 Ge Wang , Yue Wang , Michael W. Vannier

Vision-language models (VLM) have markedly advanced AI-driven interpretation and reporting of complex medical imaging, such as computed tomography (CT). Yet, existing methods largely relegate clinicians to passive observers of final…

While the field of medical image analysis has undergone a transformative shift with the integration of machine learning techniques, the main challenge of these techniques is often the scarcity of large, diverse, and well-annotated datasets.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Stefano Woerner , Arthur Jaques , Christian F. Baumgartner

Study Design: The study outlines the development of an autonomous AI system for chest X-ray (CXR) interpretation, trained on a vast dataset of over 5 million X rays sourced from healthcare systems across India. This AI system integrates…

Recent studies suggest that Visual Language Models (VLMs) hold great potential for tasks such as automated medical diagnosis. However, processing complex three-dimensional (3D) multimodal medical images poses significant challenges -…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Hao Wu , Hui Li , Yiyun Su

The development of successful artificial intelligence models for chest X-ray analysis relies on large, diverse datasets with high-quality annotations. While several databases of chest X-ray images have been released, most include disease…

Image and Video Processing · Electrical Eng. & Systems 2024-05-21 Nicolás Gaggion , Candelaria Mosquera , Lucas Mansilla , Julia Mariel Saidman , Martina Aineseder , Diego H. Milone , Enzo Ferrante

Cardiac structure segmentation from echocardiogram videos plays a crucial role in diagnosing heart disease. The combination of multi-view echocardiogram data is essential to enhance the accuracy and robustness of automated methods. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-09-21 Ziyang Zheng , Jiewen Yang , Xinpeng Ding , Xiaowei Xu , Xiaomeng Li

This study integrates PET metabolic information with CT anatomical structures to establish a 3D multimodal segmentation dataset for lymphoma based on whole-body FDG PET/CT examinations, which bridges the gap of the lack of standardised…

Image and Video Processing · Electrical Eng. & Systems 2025-12-08 Jiajun Ding , Beiyao Zhu , Xiaosheng Liu , Lishen Zhang , Zhao Liu

Background: This study proposes a Vision-Language Model (VLM) leveraging the SIGLIP encoder and Gemma-3b transformer decoder to enhance automated chronic tuberculosis (TB) screening. By integrating chest X-ray images with clinical data, the…

Medical data poses a daunting challenge for AI algorithms: it exists in many different modalities, experiences frequent distribution shifts, and suffers from a scarcity of examples and labels. Recent advances, including transformers and…