English
Related papers

Related papers: Fine-Grained Image-Text Alignment in Medical Imagi…

200 papers

Text-to-image generative models excel in creating images from text but struggle with ensuring alignment and consistency between outputs and prompts. This paper introduces TextMatch, a novel framework that leverages multimodal optimization…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Yucong Luo , Mingyue Cheng , Jie Ouyang , Xiaoyu Tao , Qi Liu

The development of AI-based methods to analyze radiology reports could lead to significant advances in medical diagnosis, from improving diagnostic accuracy to enhancing efficiency and reducing workload. However, the lack of…

Computation and Language · Computer Science 2025-08-14 Yuyan Ge , Kwan Ho Ryan Chan , Pablo Messina , René Vidal

Medical image synthesis has become an essential strategy for augmenting datasets and improving model generalization in data-scarce clinical settings. However, fine-grained and controllable synthesis remains difficult due to limited…

Image and Video Processing · Electrical Eng. & Systems 2025-09-09 Shuhan Ding , Jingjing Fu , Yu Gu , Naiteek Sangani , Mu Wei , Paul Vozila , Nan Liu , Jiang Bian , Hoifung Poon

Recently, semantic segmentation models trained with image-level text supervision have shown promising results in challenging open-world scenarios. However, these models still face difficulties in learning fine-grained semantic alignment at…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Kaixin Cai , Pengzhen Ren , Yi Zhu , Hang Xu , Jianzhuang Liu , Changlin Li , Guangrun Wang , Xiaodan Liang

Audio-text retrieval (ATR), which retrieves a relevant caption given an audio clip (A2T) and vice versa (T2A), has recently attracted much research attention. Existing methods typically aggregate information from each modality into a single…

Sound · Computer Science 2024-03-18 Qian Wang , Jia-Chen Gu , Zhen-Hua Ling

Deep neural networks excel in radiological image classification but frequently suffer from poor interpretability, limiting clinical acceptance. We present MedicalPatchNet, an inherently self-explainable architecture for chest X-ray…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Patrick Wienholt , Christiane Kuhl , Jakob Nikolas Kather , Sven Nebelung , Daniel Truhn

Multimodal medical large language models have shown substantial progress in chest X-ray interpretation but continue to face challenges in spatial reasoning and anatomical understanding. Although existing grounding techniques improve overall…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Anees Ur Rehman Hashmi , Numan Saeed , Christoph Lippert

Recent developments in the field of Natural Language Processing, especially language models such as the transformer have brought state-of-the-art results in language understanding and language generation. In this work, we investigate the…

Computation and Language · Computer Science 2024-08-22 Sonit Singh

Purpose: Interpreting chest radiographs (CXR) remains challenging due to the ambiguity of overlapping structures such as the lungs, heart, and bones. To address this issue, we propose a novel method for extracting fine-grained anatomical…

Image and Video Processing · Electrical Eng. & Systems 2023-06-08 Constantin Seibold , Alexander Jaus , Matthias A. Fink , Moon Kim , Simon Reiß , Ken Herrmann , Jens Kleesiek , Rainer Stiefelhagen

Recent works show that Generative Adversarial Networks (GANs) can be successfully applied to chest X-ray data augmentation for lung disease recognition. However, the implausible and distorted pathology features generated from the less than…

Image and Video Processing · Electrical Eng. & Systems 2020-01-23 Yunyan Xing , Zongyuan Ge , Rui Zeng , Dwarikanath Mahapatra , Jarrel Seah , Meng Law , Tom Drummond

An important component of human analysis of medical images and their context is the ability to relate newly seen things to related instances in our memory. In this paper we mimic this ability by using multi-modal retrieval augmentation and…

Computer Vision and Pattern Recognition · Computer Science 2023-02-23 Tom van Sonsbeek , Marcel Worring

Radiology Report Generation (RRG) aims to automatically generate diagnostic reports from radiology images. To achieve this, existing methods have leveraged the powerful cross-modal generation capabilities of Multimodal Large Language Models…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Jiechao Gao , Chang Liu , Yuangang Li

Radiologists highly desire fully automated versatile AI for medical imaging interpretation. However, the lack of extensively annotated large-scale multi-disease datasets has hindered the achievement of this goal. In this paper, we explore…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Weiwei Cao , Jianpeng Zhang , Yingda Xia , Tony C. W. Mok , Zi Li , Xianghua Ye , Le Lu , Jian Zheng , Yuxing Tang , Ling Zhang

We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal…

Multimedia · Computer Science 2025-07-28 Hyunwoo Oh , SeungJu Cha , Kwanyoung Lee , Si-Woo Kim , Dong-Jin Kim

Artificial intelligence (AI) shows great potential in assisting radiologists to improve the efficiency and accuracy of medical image interpretation and diagnosis. However, a versatile AI model requires large-scale data and comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Zhongyi Shui , Jianpeng Zhang , Weiwei Cao , Sinuo Wang , Ruizhe Guo , Le Lu , Lin Yang , Xianghua Ye , Tingbo Liang , Qi Zhang , Ling Zhang

Report generation models offer fine-grained textual interpretations of medical images like chest X-rays, yet they often lack interactivity (i.e. the ability to steer the generation process through user queries) and localized…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Philip Müller , Georgios Kaissis , Daniel Rueckert

Image-text matching aims to find matched cross-modal pairs accurately. While current methods often rely on projecting cross-modal features into a common embedding space, they frequently suffer from imbalanced feature representations across…

Information Retrieval · Computer Science 2024-01-19 Zuhui Wang , Yunting Yin , I. V. Ramakrishnan

Following the impressive development of LLMs, vision-language alignment in LLMs is actively being researched to enable multimodal reasoning and visual IO. This direction of research is particularly relevant to medical imaging because…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Suhyeon Lee , Won Jun Kim , Jinho Chang , Jong Chul Ye

Deep learning models for chest X-ray diagnosis are constrained by limited coverage of clinically meaningful concept combinations in publicly available training datasets. While synthetic image generation has been explored to increase data…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Amy Rafferty , Rishi Ramaesh , Ajitha Rajan

As medical imaging is central to diagnostic processes, automating the generation of radiology reports has become increasingly relevant to assist radiologists with their heavy workloads. Most current methods rely solely on global image…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Hamza Kalisch , Fabian Hörst , Jens Kleesiek , Ken Herrmann , Constantin Seibold
‹ Prev 1 3 4 5 6 7 10 Next ›