中文
相关论文

相关论文: CoCa-CXR: Contrastive Captioners Learn Strong Temp…

200 篇论文

3D captioning, which aims to describe the content of 3D scenes in natural language, remains highly challenging due to the inherent sparsity of point clouds and weak cross-modal alignment in existing methods. To address these challenges, we…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Ting Huang , Zeyu Zhang , Yemin Wang , Hao Tang

Chest X-ray report generation aims to reduce radiologists' workload by automatically producing high-quality preliminary reports. A critical yet underexplored aspect of this task is the effective use of patient-specific prior knowledge --…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Kang Liu , Zhuoqi Ma , Zikang Fang , Yunan Li , Kun Xie , Qiguang Miao

Significant methodological strides have been made toward Chest X-ray (CXR) understanding via modern vision-language models (VLMs), demonstrating impressive Visual Question Answering (VQA) and CXR report generation abilities. However,…

人工智能 · 计算机科学 2024-04-01 Seil Kang , Donghyun Kim , Junhyeok Kim , Hyo Kyung Lee , Seong Jae Hwang

We present a novel approach to Chest X-ray (CXR) Visual Question Answering (VQA), addressing both single-image image-difference questions. Single-image questions focus on abnormalities within a specific CXR ("What abnormalities are seen in…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Francesco Dalla Serra , Patrick Schrempf , Chaoyang Wang , Zaiqiao Meng , Fani Deligianni , Alison Q. O'Neil

Developing an interpretable system for generating reports in chest X-ray (CXR) analysis is becoming increasingly crucial in Computer-aided Diagnosis (CAD) systems, enabling radiologists to comprehend the decisions made by these systems.…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Trong Thang Pham , Ngoc-Vuong Ho , Nhat-Tan Bui , Thinh Phan , Patel Brijesh , Donald Adjeroh , Gianfranco Doretto , Anh Nguyen , Carol C. Wu , Hien Nguyen , Ngan Le

Self-supervised learning (SSL) has emerged as a powerful paradigm for Chest X-ray (CXR) analysis under limited annotations. Yet, existing SSL strategies remain suboptimal for medical imaging. Masked image modeling allocates substantial…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Wangyu Feng , Shawn Young , Lijian Xu

Recent artificial intelligence (AI) algorithms have achieved radiologist-level performance on various medical classification tasks. However, only a few studies addressed the localization of abnormal findings from CXR scans, which is…

图像与视频处理 · 电气工程与系统科学 2022-08-09 Hieu H. Pham , Ha Q. Nguyen , Hieu T. Nguyen , Linh T. Le , Lam Khanh

Chest X-rays or chest radiography (CXR), commonly used for medical diagnostics, typically enables limited imaging compared to computed tomography (CT) scans, which offer more detailed and accurate three-dimensional data, particularly…

图像与视频处理 · 电气工程与系统科学 2025-07-25 Noa Cahan , Eyal Klang , Galit Aviram , Yiftach Barash , Eli Konen , Raja Giryes , Hayit Greenspan

The automation of chest X-ray reporting has garnered significant interest due to the time-consuming nature of the task. However, the clinical accuracy of free-text reports has proven challenging to quantify using natural language processing…

计算机视觉与模式识别 · 计算机科学 2023-05-03 Matthias Keicher , Kamilia Zaripova , Tobias Czempiel , Kristina Mach , Ashkan Khakzar , Nassir Navab

Automated generation of clinically accurate radiology reports can improve patient care. Previous report generation methods that rely on image captioning models often generate incoherent and incorrect text due to their lack of relevant…

Recent advances in vision--language pretraining have enabled strong medical foundation models, yet most analyze radiographs in isolation, overlooking the key clinical task of comparing prior and current images to assess interval change. For…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Hanbin Ko , Kyungmin Jeon , Doowoong Choi , Chang Min Park

Multimodal alignment between language and vision is the fundamental topic in current vision-language model research. Contrastive Captioners (CoCa), as a representative method, integrates Contrastive Language-Image Pretraining (CLIP) and…

计算机视觉与模式识别 · 计算机科学 2024-01-05 Ziping Ma , Furong Xu , Jian Liu , Ming Yang , Qingpei Guo

Recently large vision-language models have shown potential when interpreting complex images and generating natural language descriptions using advanced reasoning. Medicine's inherently multimodal nature incorporating scans and text-based…

图像与视频处理 · 电气工程与系统科学 2024-07-15 Naman Sharma

Purpose: Limited studies exploring concrete methods or approaches to tackle and enhance model fairness in the radiology domain. Our proposed AI model utilizes supervised contrastive learning to minimize bias in CXR diagnosis. Materials and…

图像与视频处理 · 电气工程与系统科学 2024-01-30 Mingquan Lin , Tianhao Li , Zhaoyi Sun , Gregory Holste , Ying Ding , Fei Wang , George Shih , Yifan Peng

The chest X-ray (CXR) is commonly employed to diagnose thoracic illnesses, but the challenge of achieving accurate automatic diagnosis through this method persists due to the complex relationship between pathology. In recent years, various…

图像与视频处理 · 电气工程与系统科学 2023-05-23 Weizhi Nie , Chen Zhang , Dan song , Yunpeng Bai , Keliang Xie , Anan Liu

Following the impressive development of LLMs, vision-language alignment in LLMs is actively being researched to enable multimodal reasoning and visual IO. This direction of research is particularly relevant to medical imaging because…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Suhyeon Lee , Won Jun Kim , Jinho Chang , Jong Chul Ye

Chest X-Ray (CXR) images are commonly used for clinical screening and diagnosis. Automatically writing reports for these images can considerably lighten the workload of radiologists for summarizing descriptive findings and conclusive…

计算与语言 · 计算机科学 2020-07-24 Baoyu Jing , Zeya Wang , Eric Xing

During the COVID-19 pandemic, the sheer volume of imaging performed in an emergency setting for COVID-19 diagnosis has resulted in a wide variability of clinical CXR acquisitions. This variation is seen in the CXR projections used, image…

Vision-language models (VLMs) have shown strong promise for medical image analysis, but most remain opaque, offering predictions without the transparent, stepwise reasoning clinicians rely on. We present a framework that brings…

IMACT-CXR is an interactive multi-agent conversational tutor that helps trainees interpret chest X-rays by unifying spatial annotation, gaze analysis, knowledge retrieval, and image-grounded reasoning in a single AutoGen-based workflow. The…

人工智能 · 计算机科学 2026-04-17 Tuan-Anh Le , Anh Mai Vu , David Yang , Akash Awasthi , Hien Van Nguyen