English
Related papers

Related papers: BIMCV-R: A Landmark Dataset for 3D CT Text-Image R…

200 papers

With a widespread use of digital imaging data in hospitals, the size of medical image repositories is increasing rapidly. This causes difficulty in managing and querying these large databases leading to the need of content based medical…

Computer Vision and Pattern Recognition · Computer Science 2017-08-02 Adnan Qayyum , Syed Muhammad Anwar , Muhammad Awais , Muhammad Majid

Medical image retrieval is a valuable field for supporting clinical decision-making, yet current methods primarily support 2D images and require fully annotated queries, limiting clinical flexibility. To address this, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Inye Na , Nejung Rue , Jiwon Chung , Hyunjin Park

Automated interpretation of chest X-rays (CXR) is a critical task with the potential to significantly improve clinical workflow and patient care. While recent advances in multimodal foundation models have shown promise, effectively…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Alexander Davis , Rafael Souza , Jia-Hao Lim

This paper introduces an innovative methodology for producing high-quality 3D lung CT images guided by textual information. While diffusion-based generative models are increasingly used in medical imaging, current state-of-the-art…

Image and Video Processing · Electrical Eng. & Systems 2024-10-16 Yanwu Xu , Li Sun , Wei Peng , Shuyue Jia , Katelyn Morrison , Adam Perer , Afrooz Zandifar , Shyam Visweswaran , Motahhare Eslami , Kayhan Batmanghelich

Content-based retrieval supports a radiologist decision making process by presenting the doctor the most similar cases from the database containing both historical diagnosis and further disease development history. We present a deep…

Information Retrieval · Computer Science 2020-07-15 Ilia Kravets , Tal Heletz , Hayit Greenspan

Text-guided image editing has seen significant progress in natural image domains, but its application in medical imaging remains limited and lacks standardized evaluation frameworks. Such editing could revolutionize clinical practices by…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Minghao Liu , Zhitao He , Zhiyuan Fan , Qingyun Wang , Yi R. Fung

The amount of medical images stored in hospitals is increasing faster than ever; however, utilizing the accumulated medical images has been limited. This is because existing content-based medical image retrieval (CBMIR) systems usually…

Joint image-text embedding extracted from medical images and associated contextual reports is the bedrock for most biomedical vision-and-language (V+L) tasks, including medical visual question answering, clinical image-text retrieval,…

Computer Vision and Pattern Recognition · Computer Science 2020-09-04 Yikuan Li , Hanyin Wang , Yuan Luo

Vision-and-language models (VLMs) have been increasingly explored in the medical domain, particularly following the success of CLIP in general domain. However, unlike the relatively straightforward pairing of 2D images and text, curating…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Ziyang Zhang , Yang Yu , Xulei Yang , Si Yong Yeo

Radiologic diagnostic errors-under-reading errors, inattentional blindness, and communication failures-remain prevalent in clinical practice. These issues often stem from missed localized abnormalities, limited global context, and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Yuheng Li , Yenho Chen , Yuxiang Lai , Jike Zhong , Vanessa Wildman , Xiaofeng Yang

Content-Based Image Retrieval (CBIR) locates, retrieves and displays images alike to one given as a query, using a set of features. It demands accessible data in medical archives and from medical equipment, to infer meaning after some…

Computer Vision and Pattern Recognition · Computer Science 2016-10-11 Albany E. Herrmann , Vania Vieira Estrela

Medical images can be a valuable resource for reliable information to support medical diagnosis. However, the large volume of medical images makes it challenging to retrieve relevant information given a particular scenario. To solve this…

Computer Vision and Pattern Recognition · Computer Science 2016-10-04 S. Sharma , I. Umar , L. Ospina , D. Wong , H. R. Tizhoosh

Automated medical image analysis systems often require large amounts of training data with high quality labels, which are difficult and time consuming to generate. This paper introduces Radiology Object in COntext version 2 (ROCOv2), a…

Medical report generation is one of the most challenging tasks in medical image analysis. Although existing approaches have achieved promising results, they either require a predefined template database in order to retrieve sentences or…

Computation and Language · Computer Science 2021-06-14 Xingyi Yang , Muchao Ye , Quanzeng You , Fenglong Ma

Perceiving text is crucial to understand semantics of outdoor scenes and hence is a critical requirement to build intelligent systems for driver assistance and self-driving. Most of the existing datasets for text detection and recognition…

Computer Vision and Pattern Recognition · Computer Science 2020-05-20 Sangeeth Reddy , Minesh Mathew , Lluis Gomez , Marcal Rusinol , Dimosthenis Karatzas. , C. V. Jawahar

Semantic retrieval of remote sensing (RS) images is a critical task fundamentally challenged by the \textquote{semantic gap}, the discrepancy between a model's low-level visual features and high-level human concepts. While large…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 J. Xiao , Y. Guo , X. Zi , K. Thiyagarajan , C. Moreira , M. Prasad

Real-time magnetic resonance imaging (RT-MRI) of human speech production is enabling significant advances in speech science, linguistics, bio-inspired speech technology development, and clinical applications. Easy access to RT-MRI is…

Vision-language models (VLMs) have shown potential for automated radiology report generation, yet existing approaches rely on global embedding compression of volumetric data, often leading to hallucinated findings and limited anatomical…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Giuseppe A. Orlando , Paolo Papotti , Maria A. Zuluaga , Olivier Humbert , Marco Lorenzi

This paper presents the AToMiC (Authoring Tools for Multimedia Content) dataset, designed to advance research in image/text cross-modal retrieval. While vision-language pretrained transformers have led to significant improvements in…

We introduce TV show Retrieval (TVR), a new multimodal retrieval dataset. TVR requires systems to understand both videos and their associated subtitle (dialogue) texts, making it more realistic. The dataset contains 109K queries collected…

Computer Vision and Pattern Recognition · Computer Science 2020-08-19 Jie Lei , Licheng Yu , Tamara L. Berg , Mohit Bansal