中文
相关论文

相关论文: A Comparison of Object Detection and Phrase Ground…

200 篇论文

Unlike Object Detection, Visual Grounding task necessitates the detection of an object described by complex free-form language. To simultaneously model such complex semantic and visual representations, recent state-of-the-art studies adopt…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Weitai Kang , Luowei Zhou , Junyi Wu , Changchang Sun , Yan Yan

Phrase grounding, the problem of associating image regions to caption words, is a crucial component of vision-language tasks. We show that phrase grounding can be learned by optimizing word-region attention to maximize a lower bound on…

计算机视觉与模式识别 · 计算机科学 2020-08-07 Tanmay Gupta , Arash Vahdat , Gal Chechik , Xiaodong Yang , Jan Kautz , Derek Hoiem

In this article, we propose using deep learning and transformer architectures combined with classical machine learning algorithms to detect and identify text anomalies in texts. Deep learning model provides a very crucial context…

计算与语言 · 计算机科学 2022-11-28 Amir Jafari

This paper proposes a novel multimodal DL architecture incorporating medical images and eye-tracking data for abnormality detection in chest x-rays. Our results show that applying eye gaze data directly into DL architectures does not show…

计算机视觉与模式识别 · 计算机科学 2023-02-07 André Luís , Chihcheng Hsieh , Isabel Blanco Nobre , Sandra Costa Sousa , Anderson Maciel , Catarina Moreira , Joaquim Jorge

Instance level detection of thoracic diseases or abnormalities are crucial for automatic diagnosis in chest X-ray images. Most existing works on chest X-rays focus on disease classification and weakly supervised localization. In order to…

图像与视频处理 · 电气工程与系统科学 2020-10-20 Jingyu Liu , Jie Lian , Yizhou Yu

Fine-grained representation learning is crucial for retrieval and phrase grounding in chest X-rays, where clinically relevant findings are often spatially confined. However, the lack of region-level supervision in contrastive models and the…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Myeongkyun Kang , Yanting Yang , Xiaoxiao Li

The scarcity of richly annotated medical images is limiting supervised deep learning based solutions to medical image analysis tasks, such as localizing discriminatory radiomic disease signatures. Therefore, it is desirable to leverage…

计算机视觉与模式识别 · 计算机科学 2019-06-10 Saeid Asgari Taghanaki , Mohammad Havaei , Tess Berthier , Francis Dutil , Lisa Di Jorio , Ghassan Hamarneh , Yoshua Bengio

By analyzing human readers' performance in detecting small round lesions in simulated digital breast tomosynthesis background in a location known exactly scenario, we have developed a model observer that is a better predictor of human…

计算机视觉与模式识别 · 计算机科学 2015-07-01 Ali R. N. Avanaki , Kathryn S. Espig , Tom R. L. Kimpe , Andrew D. A. Maidment

Current approaches to explaining the decisions of deep learning systems for medical tasks have focused on visualising the elements that have contributed to each decision. We argue that such approaches are not enough to "open the black box"…

人工智能 · 计算机科学 2018-06-04 William Gale , Luke Oakden-Rayner , Gustavo Carneiro , Andrew P Bradley , Lyle J Palmer

Textual grounding, i.e., linking words to objects in images, is a challenging but important task for robotics and human-computer interaction. Existing techniques benefit from recent progress in deep learning and generally formulate the task…

计算机视觉与模式识别 · 计算机科学 2018-03-30 Raymond A. Yeh , Minh N. Do , Alexander G. Schwing

Generating accurate and clinically meaningful radiology reports from chest X-ray images remains a significant challenge in medical AI. While recent vision-language models achieve strong results in general radiology report generation, they…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Nikolay Nechaev , Evgeniia Przhezdzetskaia , Dmitry Umerenkov , Dmitry V. Dylov

Recently, chest X-ray report generation, which aims to automatically generate descriptions of given chest X-ray images, has received growing research interests. The key challenge of chest X-ray report generation is to accurately capture and…

计算机视觉与模式识别 · 计算机科学 2023-04-12 Fenglin Liu , Changchang Yin , Xian Wu , Shen Ge , Yuexian Zou , Ping Zhang , Yuexian Zou , Xu Sun

Deep learning object detection algorithm has been widely used in medical image analysis. Currently all the object detection tasks are based on the data annotated with object classes and their bounding boxes. On the other hand, medical…

计算机视觉与模式识别 · 计算机科学 2020-03-04 Li Xiao , Cheng Zhu , Junjun Liu , Chunlong Luo , Peifang Liu , Yi Zhao

Chest radiography is an effective screening tool for diagnosing pulmonary diseases. In computer-aided diagnosis, extracting the relevant region of interest, i.e., isolating the lung region of each radiography image, can be an essential step…

图像与视频处理 · 电气工程与系统科学 2022-02-23 Hilda Azimi , Jianxing Zhang , Pengcheng Xi , Hala Asad , Ashkan Ebadi , Stephane Tremblay , Alexander Wong

Chest X-ray imaging is commonly used to diagnose pneumonia, but accurately localizing the pneumonia-affected regions typically requires detailed pixel-level annotations, which are costly and time consuming to obtain. To address this…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Kiran Shahi , Anup Bagale

Creating a large-scale dataset of abnormality annotation on medical images is a labor-intensive and costly task. Leveraging weak supervision from readily available data such as radiology reports can compensate lack of large-scale data for…

计算机视觉与模式识别 · 计算机科学 2022-06-28 Ke Yu , Shantanu Ghosh , Zhexiong Liu , Christopher Deible , Kayhan Batmanghelich

Performing data augmentation for learning deep neural networks is well known to be important for training visual recognition systems. By artificially increasing the number of training examples, it helps reducing overfitting and improves…

计算机视觉与模式识别 · 计算机科学 2018-07-20 Nikita Dvornik , Julien Mairal , Cordelia Schmid

Medical language processing and deep learning techniques have emerged as critical tools for improving healthcare, particularly in the analysis of medical imaging and medical text data. These multimodal data fusion techniques help to improve…

计算与语言 · 计算机科学 2025-04-28 Sayeh Gholipour Picha , Dawood Al Chanti , Alice Caplier

Locating lesions is important in the computer-aided diagnosis of X-ray images. However, box-level annotation is time-consuming and laborious. How to locate lesions accurately with few, or even without careful annotations is an urgent…

计算机视觉与模式识别 · 计算机科学 2021-02-02 Gangming Zhao , Baolian Qi , Jinpeng Li

Visual grounding seeks to localize the image region corresponding to a free-form text description. Recently, the strong multimodal capabilities of Large Vision-Language Models (LVLMs) have driven substantial improvements in visual…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Seil Kang , Jinyeong Kim , Junhyeok Kim , Seong Jae Hwang