中文
相关论文

相关论文: Activating Associative Disease-Aware Vision Token …

200 篇论文

We introduce a radiology-focused visual language model designed to generate radiology reports from chest X-rays. Building on previous findings that large language models (LLMs) can acquire multimodal capabilities when aligned with…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Xi Zhang , Zaiqiao Meng , Jake Lever , Edmond S. L. Ho

X-ray image-based medical report generation (MRG) is a pivotal area in artificial intelligence that can significantly reduce diagnostic burdens for clinicians and patient wait times. Existing MRG models predominantly rely on Large Language…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Mingzheng Zhang , Jinfeng Gao , Dan Xu , Jiangrui Yu , Yuhan Qiao , Lan Chen , Jin Tang , Xiao Wang

Automated radiology report generation (RRG) holds potential to reduce the workload of radiologists, and recent advances in multimodal large language models (MLLMs) have enabled multimodal chest X-ray (CXR) report generation. However,…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Jonggwon Park , Byungmu Yoon , Soobum Kim , Kyoyun Choi

Vision-language models have become increasingly powerful for tasks that require an understanding of both visual and linguistic elements, bridging the gap between these modalities. In the context of multimodal clinical AI, there is a growing…

计算与语言 · 计算机科学 2024-04-30 Masoud Monajatipoor , Zi-Yi Dou , Aichi Chien , Nanyun Peng , Kai-Wei Chang

Radiology reports are detailed text descriptions of the content of medical scans. Each report describes the presence/absence and location of relevant clinical findings, commonly including comparison with prior exams of the same patient to…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Francesco Dalla Serra , Chaoyang Wang , Fani Deligianni , Jeffrey Dalton , Alison Q O'Neil

Recent advances in deep learning have enabled researchers to explore tasks at the intersection of computer vision and natural language processing, such as image captioning, visual question answering, visual dialogue, and visual language…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Sonit Singh

Artificial intelligence (AI) is reshaping modern healthcare by advancing disease diagnosis, treatment decision-making, and biomedical research. Among AI technologies, large language models (LLMs) have become especially impactful, enabling…

人工智能 · 计算机科学 2025-11-18 Zhengda Wang , Daqian Shi , Jingyi Zhao , Xiaolei Diao , Xiongfeng Tang , Yanguo Qin

Radiology report generation aims to produce computer-aided diagnoses to alleviate the workload of radiologists and has drawn increasing attention recently. However, previous deep learning methods tend to neglect the mutual influences…

计算与语言 · 计算机科学 2022-01-12 Song Wang , Liyan Tang , Mingquan Lin , George Shih , Ying Ding , Yifan Peng

Beyond generating long and topic-coherent paragraphs in traditional captioning tasks, the medical image report composition task poses more task-oriented challenges by requiring both the highly-accurate medical term diagnosis and multiple…

计算与语言 · 计算机科学 2021-01-12 Fuyu Wang , Xiaodan Liang , Lin Xu , Liang Lin

Collecting and annotating medical images is a time-consuming and resource-intensive task. However, generating synthetic data through models such as Diffusion offers a cost-effective alternative. This paper introduces a new method for the…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Ruochen Pi , Lianlei Shan

To reduce doctors' workload, deep-learning-based automatic medical report generation has recently attracted more and more research efforts, where attention mechanisms and reinforcement learning are integrated with the classic…

计算与语言 · 计算机科学 2020-11-17 Wenting Xu , Chang Qi , Zhenghua Xu , Thomas Lukasiewicz

The increasing prevalence of retinal diseases poses a significant challenge to the healthcare system, as the demand for ophthalmologists surpasses the available workforce. This imbalance creates a bottleneck in diagnosis and treatment,…

图像与视频处理 · 电气工程与系统科学 2025-09-17 Jia-Hong Huang

For robot-assisted surgery, an accurate surgical report reflects clinical operations during surgery and helps document entry tasks, post-operative analysis and follow-up treatment. It is a challenging task due to many complex and diverse…

计算机视觉与模式识别 · 计算机科学 2023-08-30 Hongqiu Wang , Yueming Jin , Lei Zhu

Medical Large Vision-Language Models (Med-LVLMs) have been widely adopted for medical report generation. Despite Med-LVLMs producing state-of-the-art performance, they exhibit a bias toward predicting all findings as normal, leading to…

多智能体系统 · 计算机科学 2025-05-27 Pengyu Wang , Shuchang Ye , Usman Naseem , Jinman Kim

Clinical practice frequently uses medical imaging for diagnosis and treatment. A significant challenge for automatic radiology report generation is that the radiology reports are long narratives consisting of multiple sentences for both…

计算与语言 · 计算机科学 2023-07-03 Kaveri Kale , pushpak Bhattacharyya , Kshitij Jadhav

Recently large vision-language models have shown potential when interpreting complex images and generating natural language descriptions using advanced reasoning. Medicine's inherently multimodal nature incorporating scans and text-based…

图像与视频处理 · 电气工程与系统科学 2024-07-15 Naman Sharma

Deep learning for histopathology has been successfully used for disease classification, image segmentation and more. However, combining image and text modalities using current state-of-the-art methods has been a challenge due to the high…

计算机视觉与模式识别 · 计算机科学 2023-11-15 Saurav Sengupta , Donald E. Brown

Medical images are widely used in clinical practice for diagnosis. Automatically generating interpretable medical reports can reduce radiologists' burden and facilitate timely care. However, most existing approaches to automatic report…

计算机视觉与模式识别 · 计算机科学 2022-11-18 Jinghan Sun , Dong Wei , Liansheng Wang , Yefeng Zheng

Reading and interpreting chest X-ray images is one of the most radiologist's routines. However, it still can be challenging, even for the most experienced ones. Therefore, we proposed a multi-model deep learning-based automated chest X-ray…

图像与视频处理 · 电气工程与系统科学 2024-01-31 Arief Purnama Muharram , Hollyana Puteri Haryono , Abassi Haji Juma , Ira Puspasari , Nugraha Priya Utama

Multimodal Large Language Models (MLLMs) have shown strong potential for radiology report generation, yet their clinical translation is hindered by architectural heterogeneity and the prevalence of factual hallucinations. Standard…

机器学习 · 计算机科学 2026-01-13 Kun Zhao , Siyuan Dai , Pan Wang , Jifeng Song , Hui Ji , Chenghua Lin , Liang Zhan , Haoteng Tang