English
Related papers

Related papers: Eyes on the Image: Gaze Supervised Multimodal Lear…

200 papers

Purpose: As visual inspection is an inherent process during radiological screening, the associated eye gaze data can provide valuable insights into relevant clinical decisions. As deep learning has become the state-of-the-art for…

Image and Video Processing · Electrical Eng. & Systems 2025-02-18 Zirui Qiu , Hassan Rivaz , Yiming Xiao

Radiology report generation from chest X-rays is an important task in artificial intelligence with the potential to greatly reduce radiologists' workload and shorten patient wait times. Despite recent advances, existing approaches often…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Puzhen Wu , Hexin Dong , Yi Lin , Yihao Ding , Yifan Peng

Image-to-text radiology report generation aims to automatically produce radiology reports that describe the findings in medical images. Most existing methods focus solely on the image data, disregarding the other patient information…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Nurbanu Aksoy , Serge Sharoff , Selcuk Baser , Nishant Ravikumar , Alejandro F Frangi

Automated radiology report generation offers an effective solution to alleviate radiologists' workload. However, most existing methods focus primarily on single or fixed-view images to model current disease conditions, which limits…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Kang Liu , Zhuoqi Ma , Xiaolu Kang , Yunan Li , Kun Xie , Zhicheng Jiao , Qiguang Miao

Automatic radiology report generation is a promising application of multimodal deep learning, aiming to reduce reporting workload and improve consistency. However, current state-of-the-art (SOTA) systems - such as Multimodal AI for…

Radiology reports are crucial for planning treatment strategies and facilitating effective doctor-patient communication. However, the manual creation of these reports places a significant burden on radiologists. While automatic radiology…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Qiguang Miao , Kang Liu , Zhuoqi Ma , Yunan Li , Xiaolu Kang , Ruixuan Liu , Tianyi Liu , Kun Xie , Zhicheng Jiao

In clinics, a radiology report is crucial for guiding a patient's treatment. However, writing radiology reports is a heavy burden for radiologists. To this end, we present an automatic, multi-modal approach for report generation from a…

Image and Video Processing · Electrical Eng. & Systems 2022-06-02 Shuxin Yang , Xian Wu , Shen Ge , S. Kevin Zhou , Li Xiao

Recent self-supervised contrastive learning methods greatly benefit from the Siamese structure that aims to minimizing distances between positive pairs. These methods usually apply random data augmentation to input images, expecting the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Sheng Wang , Zixu Zhuang , Xi Ouyang , Lichi Zhang , Zheren Li , Chong Ma , Tianming Liu , Dinggang Shen , Qian Wang

Existing deep learning methods for radiology report generation enhance diagnostic efficiency but often overlook physician-informed medical priors. This leads to a suboptimal alignment between the structured explanations and disease…

Tissues and Organs · Quantitative Biology 2026-04-13 Aishik Konwer , Moinak Bhattacharya , Prateek Prasanna

Radiology reports are detailed text descriptions of the content of medical scans. Each report describes the presence/absence and location of relevant clinical findings, commonly including comparison with prior exams of the same patient to…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Francesco Dalla Serra , Chaoyang Wang , Fani Deligianni , Jeffrey Dalton , Alison Q O'Neil

Joint embeddings between medical imaging modalities and associated radiology reports have the potential to offer significant benefits to the clinical community, ranging from cross-domain retrieval to conditional generation of reports to the…

Machine Learning · Computer Science 2018-11-28 Tzu-Ming Harry Hsu , Wei-Hung Weng , Willie Boag , Matthew McDermott , Peter Szolovits

Despite recent advances in medical vision-language pretraining, existing models still struggle to capture the diagnostic workflow: radiographs are typically treated as context-agnostic images, while radiologists' gaze -- a crucial cue for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Kang Liu , Zhuoqi Ma , Siyu Liang , Yunan Li , Xiyue Gao , Chao Liang , Kun Xie , Qiguang Miao

Obtaining large-scale radiology reports can be difficult for medical images due to various reasons, limiting the effectiveness of contrastive pre-training in the medical image domain and underscoring the need for alternative methods. In…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Zihao Zhao , Sheng Wang , Qian Wang , Dinggang Shen

Generative models have revolutionized Artificial Intelligence (AI), particularly in multimodal applications. However, adapting these models to the medical domain poses unique challenges due to the complexity of medical data and the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Daniele Molino , Francesco di Feola , Linlin Shen , Paolo Soda , Valerio Guarrasi

Visual gaze estimation, with its wide-ranging application scenarios, has garnered increasing attention within the research community. Although existing approaches infer gaze solely from image signals, recent advances in visual-language…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Jun Wang , Hao Ruan , Liangjian Wen , Yong Dai , Mingjie Wang

Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yiwei Li , Zihao Wu , Yanjun Lv , Hanqi Jiang , Weihang You , Zhengliang Liu , Dajiang Zhu , Xiang Li , Quanzheng Li , Tianming Liu , Lin Zhao

Recent medical multimodal foundation models are built as multimodal LLMs (MLLMs) by connecting a CLIP-pretrained vision encoder to an LLM using LLaVA-style finetuning. This two-stage, decoupled approach introduces a projection layer that…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Ashwin Kumar , Robbie Holland , Corey Barrett , Jangwon Kim , Maya Varma , Zhihong Chen , Yunhe Gao , Greg Zaharchuk , Tara Taghavi , Krishnaram Kenthapadi , Akshay Chaudhari

Predicting human gaze behavior within computer vision is integral for developing interactive systems that can anticipate user attention, address fundamental questions in cognitive science, and hold implications for fields like…

Image and Video Processing · Electrical Eng. & Systems 2024-07-02 Akash Awasthi , Ngan Le , Zhigang Deng , Rishi Agrawal , Carol C. Wu , Hien Van Nguyen

Multimodal deep learning utilizing imaging and diagnostic reports has made impressive progress in the field of medical imaging diagnostics, demonstrating a particularly strong capability for auxiliary diagnosis in cases where sufficient…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Hao Yang , Hong-Yu Zhou , Cheng Li , Weijian Huang , Jiarun Liu , Yong Liang , Guangming Shi , Hairong Zheng , Qiegen Liu , Shanshan Wang

Although fusion of information from multiple views of mammograms plays an important role to increase accuracy of breast cancer detection, developing multi-view mammograms-based computer-aided diagnosis (CAD) schemes still faces challenges…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Xuxin Chen , Yuheng Li , Mingzhe Hu , Ella Salari , Xiaoqian Chen , Richard L. J. Qiu , Bin Zheng , Xiaofeng Yang
‹ Prev 1 2 3 10 Next ›