中文
相关论文

相关论文: RadZero: Similarity-Based Cross-Attention for Expl…

200 篇论文

Recent advancements in Computer Assisted Diagnosis have shown promising performance in medical imaging tasks, particularly in chest X-ray analysis. However, the interaction between these models and radiologists has been primarily limited to…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Yunsoo Kim , Jinge Wu , Yusuf Abdulle , Yue Gao , Honghan Wu

Automated interpretation of chest X-rays (CXR) is a critical task with the potential to significantly improve clinical workflow and patient care. While recent advances in multimodal foundation models have shown promise, effectively…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Alexander Davis , Rafael Souza , Jia-Hao Lim

Foundation models, trained on vast amounts of data using self-supervised techniques, have emerged as a promising frontier for advancing artificial intelligence (AI) applications in medicine. This study evaluates three different…

Vision-language pre-training for chest X-rays has made significant strides, primarily by utilizing paired radiographs and radiology reports. However, existing approaches often face challenges in encoding medical knowledge effectively. While…

计算机视觉与模式识别 · 计算机科学 2024-04-05 Haozhe Luo , Ziyu Zhou , Corentin Royer , Anjany Sekuboyina , Bjoern Menze

The Critical View of Safety (CVS) is crucial for safe laparoscopic cholecystectomy, yet assessing CVS criteria remains a complex and challenging task, even for experts. Traditional models for CVS recognition depend on vision-only models…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Britty Baby , Vinkle Srivastav , Pooja P. Jain , Kun Yuan , Pietro Mascagni , Nicolas Padoy

Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Yang Liu , Mengyuan Liu , Shudong Huang , Jiancheng Lv

The rapid evolution of artificial intelligence, especially in large language models (LLMs), has significantly impacted various domains, including healthcare. In chest X-ray (CXR) analysis, previous studies have employed LLMs, but with…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Jonggwon Park , Soobum Kim , Byungmu Yoon , Jihun Hyun , Kyoyun Choi

This paper proposes a novel framework for lung segmentation in chest X-rays. It consists of two key contributions, a criss-cross attention based segmentation network and radiorealistic chest X-ray image synthesis (i.e. a synthesized…

计算机视觉与模式识别 · 计算机科学 2019-04-22 Youbao Tang , Yuxing Tang , Jing Xiao , Ronald M. Summers

Medical image analysis using computer-based algorithms has attracted considerable attention from the research community and achieved tremendous progress in the last decade. With recent advances in computing resources and availability of…

图像与视频处理 · 电气工程与系统科学 2023-10-03 Huyen Tran , Duc Thanh Nguyen , John Yearwood

Medical image segmentation has significantly benefitted thanks to deep learning architectures. Furthermore, semi-supervised learning (SSL) has recently been a growing trend for improving a model's overall performance by leveraging abundant…

图像与视频处理 · 电气工程与系统科学 2021-10-05 S. M. Kamrul Hasan , Cristian A. Linte

Previous foundation models for fundus images were pre-trained with limited disease categories and knowledge base. Here we introduce a knowledge-rich vision-language model (RetiZero) that leverages knowledge from more than 400 fundus…

Semantic segmentation is a crucial task in computer vision that involves segmenting images into semantically meaningful regions at the pixel level. However, existing approaches often rely on expensive human annotations as supervision for…

计算机视觉与模式识别 · 计算机科学 2023-06-02 Jun Chen , Deyao Zhu , Guocheng Qian , Bernard Ghanem , Zhicheng Yan , Chenchen Zhu , Fanyi Xiao , Mohamed Elhoseiny , Sean Chang Culatana

Lip Reading, or Visual Automatic Speech Recognition (V-ASR), is a complex task requiring the interpretation of spoken language exclusively from visual cues, primarily lip movements and facial expressions. This task is especially challenging…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Marshall Thomas , Edward Fish , Richard Bowden

Recent advancements in multimodal Large Language Models (LLMs) have significantly enhanced the automation of medical image analysis, particularly in generating radiology reports from chest X-rays (CXR). However, these models still suffer…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Yunsoo Kim , Jinge Wu , Su-Hwan Kim , Pardeep Vasudev , Jiashu Shen , Honghan Wu

Analyzing radiology reports is a time-consuming and error-prone task, which raises the need for an efficient automated radiology report analysis system to alleviate the workloads of radiologists and encourage precise diagnosis. In this…

计算与语言 · 计算机科学 2022-04-21 Song Wang , Mingquan Lin , Ying Ding , George Shih , Zhiyong Lu , Yifan Peng

Medical eye-tracking data is an important information source for understanding how radiologists visually interpret medical images. This information not only improves the accuracy of deep learning models for X-ray analysis but also their…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Trong Thang Pham , Tien-Phat Nguyen , Yuki Ikebe , Akash Awasthi , Zhigang Deng , Carol C. Wu , Hien Nguyen , Ngan Le

While Multi-Task Learning (MTL) offers inherent advantages in complex domains such as medical imaging by enabling shared representation learning, effectively balancing task contributions remains a significant challenge. This paper addresses…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Youssef Mohamed , Noran Mohamed , Khaled Abouhashad , Feilong Tang , Sara Atito , Shoaib Jameel , Imran Razzak , Ahmed B. Zaky

We explore a novel zero-shot Audio-Visual Speech Recognition (AVSR) framework, dubbed Zero-AVSR, which enables speech recognition in target languages without requiring any audio-visual speech data in those languages. Specifically, we…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Jeong Hun Yeo , Minsu Kim , Chae Won Kim , Stavros Petridis , Yong Man Ro

In zero-shot image recognition tasks, humans demonstrate remarkable flexibility in classifying unseen categories by composing known simpler concepts. However, existing vision-language models (VLMs), despite achieving significant progress…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Hui Liu , Wenya Wang , Kecheng Chen , Jie Liu , Yibing Liu , Tiexin Qin , Peisong He , Xinghao Jiang , Haoliang Li

Recently a number of studies demonstrated impressive performance on diverse vision-language multi-modal tasks such as image captioning and visual question answering by extending the BERT architecture with multi-modal pre-training…

计算机视觉与模式识别 · 计算机科学 2022-09-22 Jong Hak Moon , Hyungyung Lee , Woncheol Shin , Young-Hak Kim , Edward Choi