中文
相关论文

相关论文: Intensive Vision-guided Network for Radiology Repo…

200 篇论文

Large language models (LLMs) have demonstrated remarkable capabilities in various domains, including radiology report generation. Previous approaches have attempted to utilize multimodal LLMs for this task, enhancing their performance…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Wenjun Hou , Yi Cheng , Kaishuai Xu , Heng Li , Yan Hu , Wenjie Li , Jiang Liu

Automated pathology report generation from Whole Slide Images (WSIs) faces two key challenges: (1) lack of semantic content in visual features and (2) inherent information redundancy in WSIs. To address these issues, we propose a novel…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Ling Zhang , Boxiang Yun , Qingli Li , Yan Wang

Generating long and coherent reports to describe medical images poses challenges to bridging visual patterns with informative human linguistic descriptions. We propose a novel Hybrid Retrieval-Generation Reinforced Agent (HRGR-Agent) which…

计算机视觉与模式识别 · 计算机科学 2018-11-27 Christy Y. Li , Xiaodan Liang , Zhiting Hu , Eric P. Xing

Radiology Report Generation (RRG) is essential for computer-aided diagnosis and medication guidance, which can relieve the heavy burden of radiologists by automatically generating the corresponding radiology reports according to the given…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Weixing Chen , Yang Liu , Ce Wang , Jiarui Zhu , Guanbin Li , Cheng-Lin Liu , Liang Lin

In response to the worldwide COVID-19 pandemic, advanced automated technologies have emerged as valuable tools to aid healthcare professionals in managing an increased workload by improving radiology report generation and prognostic…

图像与视频处理 · 电气工程与系统科学 2024-05-24 Zhusi Zhong , Jie Li , John Sollee , Scott Collins , Harrison Bai , Paul Zhang , Terrence Healey , Michael Atalay , Xinbo Gao , Zhicheng Jiao

We propose Retrieval Augmented Generation (RAG) as an approach for automated radiology report writing that leverages multimodally aligned embeddings from a contrastively pretrained vision language model for retrieval of relevant candidate…

计算与语言 · 计算机科学 2023-05-08 Mercy Ranjit , Gopinath Ganapathy , Ranjit Manuel , Tanuja Ganu

Chest X-ray is one of the most popular medical imaging modalities due to its accessibility and effectiveness. However, there is a chronic shortage of well-trained radiologists who can interpret these images and diagnose the patient's…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Otabek Nazarov , Mohammad Yaqub , Karthik Nandakumar

Automated radiology report generation aims to generate radiology reports that contain rich, fine-grained descriptions of radiology imaging. Compared with image captioning in the natural image domain, medical images are very similar to each…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Yuhao Wang

Generating images according to natural language descriptions is a challenging task. Prior research has mainly focused to enhance the quality of generation by investigating the use of spatial attention and/or textual attention thereby…

计算机视觉与模式识别 · 计算机科学 2022-01-17 Henning Schulze , Dogucan Yaman , Alexander Waibel

This paper explores the task of radiology report generation, which aims at generating free-text descriptions for a set of radiographs. One significant challenge of this task is how to correctly maintain the consistency between the images…

计算与语言 · 计算机科学 2023-06-13 Wenjun Hou , Kaishuai Xu , Yi Cheng , Wenjie Li , Jiang Liu

Accurate yet interpretable image-based diagnosis remains a central challenge in medical AI, particularly in settings characterized by limited data, subtle visual cues, and high-stakes clinical decision-making. Most existing vision models…

人工智能 · 计算机科学 2025-12-23 Midhat Urooj , Ayan Banerjee , Sandeep Gupta

We propose a learned image-guided rendering technique that combines the benefits of image-based rendering and GAN-based image synthesis. The goal of our method is to generate photo-realistic re-renderings of reconstructed objects for…

计算机视觉与模式识别 · 计算机科学 2020-01-16 Justus Thies , Michael Zollhöfer , Christian Theobalt , Marc Stamminger , Matthias Nießner

VQA (Visual Question Answering) combines Natural Language Processing (NLP) with image understanding to answer questions about a given image. It has enormous potential for the development of medical diagnostic AI systems. Such a system can…

图像与视频处理 · 电气工程与系统科学 2025-07-30 Gaurav Parajuli

Visual question answering (VQA) in medical imaging aims to support clinical diagnosis by automatically interpreting complex imaging data in response to natural language queries. Existing studies typically rely on distinct visual and textual…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Yuanhe Tian , Chen Su , Junwen Duan , Yan Song

Graphic design visually conveys information and data by creating and combining text, images and graphics. Two-stage methods that rely primarily on layout generation lack creativity and intelligence, making graphic design still…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Yadong Qu , Shancheng Fang , Yuxin Wang , Xiaorui Wang , Zhineng Chen , Hongtao Xie , Yongdong Zhang

Automatic generation of radiology reports seeks to reduce clinician workload while improving documentation consistency. Existing methods that adopt encoder-decoder or retrieval-augmented pipelines achieve progress in fluency but remain…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Rong Fu , Yiqing Lyu , Chunlei Meng , Muge Qi , Yabin Jin , Qi Zhao , Li Bao , Juntao Gao , Fuqian Shi , Nilanjan Dey , Wei Luo , Simon Fong

Radiology report generation aims at generating descriptive text from radiology images automatically, which may present an opportunity to improve radiology reporting and interpretation. A typical setting consists of training encoder-decoder…

计算与语言 · 计算机科学 2021-09-28 An Yan , Zexue He , Xing Lu , Jiang Du , Eric Chang , Amilcare Gentili , Julian McAuley , Chun-Nan Hsu

The rapid proliferation of AI-Generated Images (AIGIs) has introduced severe risks of misinformation, making AIGI detection a critical yet challenging task. While traditional detection paradigms mainly rely on low-level features, recent…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Chenyang Zhu , Maorong Wang , Jun Liu , Ching-Chun Chang , Isao Echizen

Retrieval augmented generation (RAG) has transformed text based question answering, yet its extension to visual domains remains hindered by fundamental challenges: bridging the modality gap between image queries and text heavy knowledge…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Parthaw Goswami , Jaynto Goswami Deep

The growing realism of AI-generated images produced by recent GAN and diffusion models has intensified concerns over the reliability of visual media. Yet, despite notable progress in deepfake detection, current forensic systems degrade…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Anshul Bagaria