English
Related papers

Related papers: DiA-gnostic VLVAE: Disentangled Alignment-Constrai…

200 papers

Reliable deployment of Vision-Language Models (VLMs) in radiology requires validation metrics that go beyond surface-level text similarity to ensure clinical fidelity and demographic fairness. This paper investigates a critical blind spot…

Computation and Language · Computer Science 2026-03-03 Aditya Parikh , Aasa Feragen , Sneha Das , Stella Frank

Large vision language models (VLMs) have progressed incredibly from research to applicability for general-purpose use cases. LLaVA-Med, a pioneering large language and vision assistant for biomedicine, can perform multi-modal biomedical…

Automating medical reports for retinal images requires a sophisticated blend of visual pattern recognition and deep clinical knowledge. Current Large Vision-Language Models (LVLMs) often struggle in specialized medical fields where data is…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Nagur Shareef Shaik , Teja Krishna Cherukuri , Dong Hye Ye

We view disentanglement learning as discovering an underlying structure that equivariantly reflects the factorized variations shown in data. Traditionally, such a structure is fixed to be a vector space with data variations represented by…

Machine Learning · Computer Science 2021-06-08 Xinqi Zhu , Chang Xu , Dacheng Tao

Ophthalmologists typically require multimodal data sources to improve diagnostic accuracy in clinical decisions. However, due to medical device shortages, low-quality data and data privacy concerns, missing data modalities are common in…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Chengzhi Liu , Zile Huang , Zhe Chen , Feilong Tang , Yu Tian , Zhongxing Xu , Zihong Luo , Yalin Zheng , Yanda Meng

This paper explores training medical vision-language models (VLMs) -- where the visual and language inputs are embedded into a common space -- with a particular focus on scenarios where training data is limited, as is often the case in…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Rhydian Windsor , Amir Jamaludin , Timor Kadir , Andrew Zisserman

Most existing CLIP-style medical vision--language pretraining methods rely on global or local alignment with substantial paired data. However, global alignment is easily dominated by non-diagnostic information, while local alignment fails…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Huimin Yan , Liang Bai , Xian Yang , Long Chen

We propose a deep mixture of multimodal hierarchical variational auto-encoders called MMHVAE that synthesizes missing images from observed images in different modalities. MMHVAE's design focuses on tackling four challenges: (i) creating a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Reuben Dorent , Nazim Haouchine , Alexandra Golby , Sarah Frisken , Tina Kapur , William Wells

Existing Vision-Language Models (VLMs) are predominantly trained on web-scraped, noisy image-text data, exhibiting limited exposure to the specialized domain of RS. This deficiency results in poor performance on RS-specific tasks, as…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Angelos Zavras , Dimitrios Michail , Xiao Xiang Zhu , Begüm Demir , Ioannis Papoutsis

Radiology Report Generation (RRG) aims to produce accurate and coherent diagnostics from medical images. Although large vision language models (LVLM) improve report fluency and accuracy, they exhibit hallucinations, generating plausible yet…

Computation and Language · Computer Science 2026-02-05 Ruixiao Yang , Yuanhe Tian , Xu Yang , Huiqi Li , Yan Song

Cross-modal retrieval is crucial in understanding latent correspondences across modalities. However, existing methods implicitly assume well-matched training data, which is impractical as real-world data inevitably involves imperfect…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Zhuohang Dang , Minnan Luo , Jihong Wang , Chengyou Jia , Haochen Han , Herun Wan , Guang Dai , Xiaojun Chang , Jingdong Wang

Visual Question Answering (VQA) becomes one of the most active research problems in the medical imaging domain. A well-known VQA challenge is the intrinsic diversity between the image and text modalities, and in the medical VQA task, there…

Computer Vision and Pattern Recognition · Computer Science 2023-02-28 Yuan Zhou , Jing Mei , Yiqin Yu , Tanveer Syeda-Mahmood

Multimodal Sentiment Analysis (MSA) leverages heterogeneous modalities, such as language, vision, and audio, to enhance the understanding of human sentiment. While existing models often focus on extracting shared information across…

Machine Learning · Computer Science 2025-04-10 Pan Wang , Qiang Zhou , Yawen Wu , Tianlong Chen , Jingtong Hu

Disentangled representation learning aims to learn low-dimensional representations where each dimension corresponds to an underlying generative factor. While the Variational Auto-Encoder (VAE) is widely used for this purpose, most existing…

Machine Learning · Computer Science 2024-12-31 Di Fan , Yannian Kou , Chuanhou Gao

Diffusion probabilistic models (DPMs) have shown remarkable results on various image synthesis tasks such as text-to-image generation and image inpainting. However, compared to other generative methods like VAEs and GANs, DPMs lack a…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Yipeng Leng , Qiangjuan Huang , Zhiyuan Wang , Yangyang Liu , Haoyu Zhang

Vision-Language Models (VLMs) have significantly advanced automated Radiology Report Generation (RRG). However, existing methods implicitly assume high-quality inputs, overlooking the noise and artifacts prevalent in real-world clinical…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Hongze Zhu , Chen Hu , Jiaxuan Jiang , Hong Liu , Yawen Huang , Ming Hu , Tianyu Wang , Zhijian Wu , Yefeng Zheng

Learning robust representations of authorial style is crucial for authorship attribution and AI-generated text detection. However, existing methods often struggle with content-style entanglement, where models learn spurious correlations…

Computation and Language · Computer Science 2026-04-24 Hieu Man , Van-Cuong Pham , Nghia Trung Ngo , Franck Dernoncourt , Thien Huu Nguyen

Recent advances in AI combine large language models (LLMs) with vision encoders that bring forward unprecedented technical capabilities to leverage for a wide range of healthcare applications. Focusing on the domain of radiology,…

Oversight AI is an emerging concept in radiology where the AI forms a symbiosis with radiologists by continuously supporting radiologists in their decision-making. Recent advances in vision-language models sheds a light on the long-standing…

Image and Video Processing · Electrical Eng. & Systems 2023-04-13 Sangjoon Park , Eun Sun Lee , Kyung Sook Shin , Jeong Eun Lee , Jong Chul Ye

Recently, Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in visual understanding and reasoning across various vision-language tasks. However, we found that MLLMs cannot process effectively from…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Bangyan Li , Wenxuan Huang , Zhenkun Gao , Yeqiang Wang , Yunhang Shen , Jingzhong Lin , Ling You , Yuxiang Shen , Shaohui Lin , Wanli Ouyang , Yuling Sun