中文
相关论文

相关论文: Cross-Modal Causal Intervention for Medical Report…

200 篇论文

Causal Disentangled Representation Learning(CDRL) aims to learn and disentangle low dimensional representations and their underlying causal structure from observations. However, existing disentanglement methods rely on a standard mean-field…

机器学习 · 计算机科学 2026-01-30 Yutao Jin , Yuang Tao , Junyong Zhai

Despite significant advancements in adapting Large Language Models (LLMs) for radiology report generation (RRG), clinical adoption remains challenging due to difficulties in accurately mapping pathological and anatomical features to their…

计算机视觉与模式识别 · 计算机科学 2025-08-22 Qilong Xing , Zikai Song , Youjia Zhang , Na Feng , Junqing Yu , Wei Yang

Radiology Report Generation (RRG) aims to produce accurate and coherent diagnostics from medical images. Although large vision language models (LVLM) improve report fluency and accuracy, they exhibit hallucinations, generating plausible yet…

计算与语言 · 计算机科学 2026-02-05 Ruixiao Yang , Yuanhe Tian , Xu Yang , Huiqi Li , Yan Song

Automating radiology report generation can significantly reduce the workload of radiologists and enhance the accuracy, consistency, and efficiency of clinical documentation.We propose a novel cross-modal framework that uses MedCLIP as both…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Qianhao Han , Junyi Liu , Zengchang Qin , Zheng Zheng

The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-assisted reporting. While existing vision-language models (VLMs) show promise in specific tasks…

Medical video diagnosis involves inferring clinical decisions from dynamic tissue responses throughout examination processes. Existing methods rely on an end-to-end learning paradigm that i) focuses on appearance rather than pathology, ii)…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Jianzhe Gao , Churan Wang , Weiyi Zhang , Jianghua Li , Li-An Li , Wenguan Wang , Yixin Zhu , Yizhou Wang

Vision-language pretraining has been shown to produce high-quality visual encoders which transfer efficiently to downstream computer vision tasks. Contrastive learning approaches have increasingly been adopted for medical vision language…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Keegan Quigley , Miriam Cha , Josh Barua , Geeticka Chauhan , Seth Berkowitz , Steven Horng , Polina Golland

Automatic radiology report generation is a promising application of multimodal deep learning, aiming to reduce reporting workload and improve consistency. However, current state-of-the-art (SOTA) systems - such as Multimodal AI for…

Causal Representation Learning (CRL) aims at identifying high-level causal factors and their relationships from high-dimensional observations, e.g., images. While most CRL works focus on learning causal representations in a single…

机器学习 · 计算机科学 2024-03-18 Davide Talon , Phillip Lippe , Stuart James , Alessio Del Bue , Sara Magliacane

Radiology report generation (RRG) is commonly formulated as a single-path generation task, where a multimodal large language model (MLLM) produces one decoded report as the final output. While recent progress has largely been driven by…

计算与语言 · 计算机科学 2026-05-29 Xi Zhang , Yingshu Li , Zaiqiao Meng , Jake Lever , Edmond S. L. Ho

Multimodal large language models (MLLMs) have recently achieved remarkable progress in radiology by integrating visual perception with natural language understanding. However, they often generate clinically unsupported descriptions, known…

计算与语言 · 计算机科学 2025-10-20 Xi Zhang , Zaiqiao Meng , Jake Lever , Edmond S. L. Ho

Retrieval-Augmented Generation (RAG) has become a core paradigm in document question answering tasks. However, existing methods have limitations when dealing with multimodal documents: one category of methods relies on layout analysis and…

计算与语言 · 计算机科学 2026-03-09 Wang Chen , Wenhan Yu , Guanqiang Qi , Weikang Li , Yang Li , Lei Sha , Deguo Xia , Jizhou Huang

This thesis develops methods for causal inference and causal representation learning (CRL) in high-dimensional, time-varying data. The first contribution introduces the Causal Dynamic Variational Autoencoder (CDVAE), a model for estimating…

机器学习 · 统计学 2025-12-05 Mouad EL Bouchattaoui

Causal reasoning can be considered a cornerstone of intelligent systems. Having access to an underlying causal graph comes with the promise of cause-effect estimation and the identification of efficient and safe interventions. However,…

Causal representation learning (CRL) has garnered increasing interest from the causal inference and artificial intelligence communities due to its potential to disentangle complex data-generating mechanism into causally interpretable latent…

机器学习 · 统计学 2026-05-28 Hao Chen , Lin Liu , Yu Guang Wang

Decoding neural visual representations from electroencephalogram (EEG)-based brain activity is crucial for advancing brain-machine interfaces (BMI) and has transformative potential for neural sensory rehabilitation. While multimodal…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Yueyang Li , Zijian Kang , Shengyu Gong , Wenhao Dong , Weiming Zeng , Hongjie Yan , Wai Ting Siok , Nizhuan Wang

Long-term action recognition (LTAR) is challenging due to extended temporal spans with complex atomic action correlations and visual confounders. Although vision-language models (VLMs) have shown promise, they often rely on statistical…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Xu Shaowu , Jia Xibin , Gao Junyu , Sun Qianmei , Chang Jing , Fan Chao

We present MMCORE, a unified framework designed for multimodal image generation and editing. MMCORE leverages a pre-trained Vision-Language Model (VLM) to predict semantic visual embeddings via learnable query tokens, which subsequently…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Zijie Li , Yichun Shi , Jingxiang Sun , Ye Wang , Yixuan Huang , Zhiyao Guo , Xiaochen Lian , Peihao Zhu , Yu Tian , Zhonghua Zhai , Peng Wang

Medical image segmentation is more clinically valuable when it supports diagnosis rather than merely producing lesion masks. However, diagnostically relevant lesion cues are often subtle and localized, while existing models may be…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Fengyi Zhang , Xujie Zeng , Mohan Liu , Zengyi Wang , Yalong Jiang

Automated radiology report generation (RRG) holds potential to reduce the workload of radiologists, and recent advances in multimodal large language models (MLLMs) have enabled multimodal chest X-ray (CXR) report generation. However,…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Jonggwon Park , Byungmu Yoon , Soobum Kim , Kyoyun Choi