中文
相关论文

相关论文: Seeing Through Experts Eyes A Foundational Vision …

200 篇论文

Chest X-rays (CXRs) are among the most frequently performed imaging examinations worldwide, yet rising imaging volumes increase radiologist workload and the risk of diagnostic errors. Although artificial intelligence (AI) systems have shown…

Despite recent advances in medical vision-language pretraining, existing models still struggle to capture the diagnostic workflow: radiographs are typically treated as context-agnostic images, while radiologists' gaze -- a crucial cue for…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Kang Liu , Zhuoqi Ma , Siyu Liang , Yunan Li , Xiyue Gao , Chao Liang , Kun Xie , Qiguang Miao

Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading; yet, existing methods exploit this signal only partially, either as a static spatial prior or as an auxiliary…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Yiwei Li , Zihao Wu , Huaqin Zhao , Yifan Zhou , Chao Cao , Dajiang Zhu , Tianming Liu , Lin Zhao

Radiologists rely on eye movements to navigate and interpret medical images. A trained radiologist possesses knowledge about the potential diseases that may be present in the images and, when searching, follows a mental checklist to locate…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Trong-Thang Pham , Anh Nguyen , Zhigang Deng , Carol C. Wu , Hien Van Nguyen , Ngan Le

Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Yiwei Li , Zihao Wu , Yanjun Lv , Hanqi Jiang , Weihang You , Zhengliang Liu , Dajiang Zhu , Xiang Li , Quanzheng Li , Tianming Liu , Lin Zhao

Vision-language models (VLMs) have shown strong promise for medical image analysis, but most remain opaque, offering predictions without the transparent, stepwise reasoning clinicians rely on. We present a framework that brings…

Large Vision-Language Models (LVLMs) have demonstrated promising performance in chest X-ray (CXR) analysis. To enhance human-computer interaction, several studies have incorporated radiologists' eye gaze, typically through heatmaps or…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Yunsoo Kim , Jinge Wu , Honghan Wu

Existing deep learning methods for radiology report generation enhance diagnostic efficiency but often overlook physician-informed medical priors. This leads to a suboptimal alignment between the structured explanations and disease…

组织与器官 · 定量生物学 2026-04-13 Aishik Konwer , Moinak Bhattacharya , Prateek Prasanna

[18F]FDG-PET/CT is a cornerstone imaging modality for tumor staging and treatment response assessment across many cancer types, yet expert reader shortages necessitate more efficient diagnostic aids. While standalone AI models for automatic…

Recent advances in reasoning-enhanced large language models (LLMs) and multimodal LLMs (MLLMs) have significantly improved performance in complex tasks, yet medical AI models often overlook the structured reasoning processes inherent in…

人工智能 · 计算机科学 2025-05-22 Ziqing Fan , Cheng Liang , Chaoyi Wu , Ya Zhang , Yanfeng Wang , Weidi Xie

Automated chest radiographs interpretation requires both accurate disease classification and detailed radiology report generation, presenting a significant challenge in the clinical workflow. Current approaches either focus on…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Difei Gu , Yunhe Gao , Yang Zhou , Mu Zhou , Dimitris Metaxas

While exploring visual scenes, humans' scanpaths are driven by their underlying attention processes. Understanding visual scanpaths is essential for various applications. Traditional scanpath models predict the where and when of gaze shifts…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Xianyu Chen , Ming Jiang , Qi Zhao

Recent advancements in Computer Assisted Diagnosis have shown promising performance in medical imaging tasks, particularly in chest X-ray analysis. However, the interaction between these models and radiologists has been primarily limited to…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Yunsoo Kim , Jinge Wu , Yusuf Abdulle , Yue Gao , Honghan Wu

Developing an interpretable system for generating reports in chest X-ray (CXR) analysis is becoming increasingly crucial in Computer-aided Diagnosis (CAD) systems, enabling radiologists to comprehend the decisions made by these systems.…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Trong Thang Pham , Ngoc-Vuong Ho , Nhat-Tan Bui , Thinh Phan , Patel Brijesh , Donald Adjeroh , Gianfranco Doretto , Anh Nguyen , Carol C. Wu , Hien Nguyen , Ngan Le

In the field of chest X-ray (CXR) diagnosis, existing works often focus solely on determining where a radiologist looks, typically through tasks such as detection, segmentation, or classification. However, these approaches are often…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Trong Thang Pham , Jacob Brecheisen , Anh Nguyen , Hien Nguyen , Ngan Le

Gaze estimation is pivotal in human scene comprehension tasks, particularly in medical diagnostic analysis. Eye-tracking technology facilitates the recording of physicians' ocular movements during image interpretation, thereby elucidating…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Shaonan Liu , Wenting Chen , Jie Liu , Xiaoling Luo , Linlin Shen

Medical eye-tracking data is an important information source for understanding how radiologists visually interpret medical images. This information not only improves the accuracy of deep learning models for X-ray analysis but also their…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Trong Thang Pham , Tien-Phat Nguyen , Yuki Ikebe , Akash Awasthi , Zhigang Deng , Carol C. Wu , Hien Nguyen , Ngan Le

Large Vision Language Models (LVLMs) show immense potential for automated ophthalmic diagnosis. However, their clinical deployment is severely hindered by lacking domain-specific knowledge. In this work, we identify two structural…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Shuai Lu , Meng Wang , Jia Guo , Jiawei Du , Bo Liu , Shengzhu Yang , Weihang Zhang , Huazhu Fu , Huiqi Li

We present ReXVQA, the largest and most comprehensive benchmark for visual question answering (VQA) in chest radiology, comprising approximately 696,000 questions paired with 160,000 chest X-rays studies across training, validation, and…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Ankit Pal , Jung-Oh Lee , Xiaoman Zhang , Malaikannan Sankarasubbu , Seunghyeon Roh , Won Jung Kim , Meesun Lee , Pranav Rajpurkar

We introduce MATEX (Multi-scale Attention and Text-guided Explainability), a novel framework that advances interpretability in medical vision-language models by incorporating anatomically informed spatial reasoning. MATEX synergistically…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Muhammad Imran , Chi Lee , Yugyung Lee
‹ 上一页 1 2 3 10 下一页 ›