中文
相关论文

相关论文: A Comparison of Object Detection and Phrase Ground…

200 篇论文

Multimodal reference resolution, including phrase grounding, aims to understand the semantic relations between mentions and real-world objects. Phrase grounding between images and their captions is a well-established task. In contrast, for…

计算与语言 · 计算机科学 2025-06-03 Shun Inadumi , Nobuhiro Ueda , Koichiro Yoshino

Recent advances in deep learning have led to a promising performance in many medical image analysis tasks. As the most commonly performed radiological exam, chest radiographs are a particularly important modality for which a variety of…

图像与视频处理 · 电气工程与系统科学 2021-06-08 Ecem Sogancioglu , Erdi Çallı , Bram van Ginneken , Kicky G. van Leeuwen , Keelin Murphy

Medical phrase grounding (MPG) aims to locate the most relevant region in a medical image, given a phrase query describing certain medical findings, which is an important task for medical image analysis and radiological diagnosis. However,…

计算机视觉与模式识别 · 计算机科学 2023-03-15 Zhihao Chen , Yang Zhou , Anh Tran , Junting Zhao , Liang Wan , Gideon Ooi , Lionel Cheng , Choon Hua Thng , Xinxing Xu , Yong Liu , Huazhu Fu

In this thesis, we study multiple tasks related to document layout analysis such as the detection of text lines, the splitting into acts or the detection of the writing support. Thus, we propose two deep neural models following two…

计算机视觉与模式识别 · 计算机科学 2023-01-30 Mélodie Boillet

Self-supervised learning provides an opportunity to explore unlabeled chest X-rays and their associated free-text reports accumulated in clinical routine without manual supervision. This paper proposes a Joint Image Text Representation…

机器学习 · 计算机科学 2021-09-07 Zhanghexuan Ji , Mohammad Abuzar Shaikh , Dana Moukheiber , Sargur Srihari , Yifan Peng , Mingchen Gao

Chest X-ray scan is a most often used modality by radiologists to diagnose many chest related diseases in their initial stages. The proposed system aids the radiologists in making decision about the diseases found in the scans more…

图像与视频处理 · 电气工程与系统科学 2020-08-07 Ahmed Rasheed , Muhammad Shahzad Younis , Muhammad Bilal , Maha Rasheed

Pathology detection and delineation enables the automatic interpretation of medical scans such as chest X-rays while providing a high level of explainability to support radiologists in making informed decisions. However, annotating…

计算机视觉与模式识别 · 计算机科学 2023-09-07 Philip Müller , Felix Meissen , Johannes Brandt , Georgios Kaissis , Daniel Rueckert

Machine learning models for radiology benefit from large-scale data sets with high quality labels for abnormalities. We curated and analyzed a chest computed tomography (CT) data set of 36,316 volumes from 19,993 unique patients. This is…

图像与视频处理 · 电气工程与系统科学 2020-10-14 Rachel Lea Draelos , David Dov , Maciej A. Mazurowski , Joseph Y. Lo , Ricardo Henao , Geoffrey D. Rubin , Lawrence Carin

Finding diseases from an X-ray image is an important yet highly challenging task. Current methods for solving this task exploit various characteristics of the chest X-ray image, but one of the most important characteristics is still…

计算机视觉与模式识别 · 计算机科学 2020-07-16 Minchul Kim , Jongchan Park , Seil Na , Chang Min Park , Donggeun Yoo

We propose a method to improve Visual Question Answering (VQA) with Retrieval-Augmented Generation (RAG) by introducing text-grounded object localization. Rather than retrieving information based on the entire image, our approach enables…

人工智能 · 计算机科学 2025-10-01 Xinxi Chen , Tianyang Chen , Lijia Hong

Given a textual phrase and an image, the visual grounding problem is the task of locating the content of the image referenced by the sentence. It is a challenging task that has several real-world applications in human-computer interaction,…

计算机视觉与模式识别 · 计算机科学 2022-02-03 Davide Rigoni , Luciano Serafini , Alessandro Sperduti

Obtaining automated preliminary read reports for common exams such as chest X-rays will expedite clinical workflows and improve operational efficiencies in hospitals. However, the quality of reports generated by current automated approaches…

Using picture description speech for dementia detection has been studied for 30 years. Despite the long history, previous models focus on identifying the differences in speech patterns between healthy subjects and patients with dementia but…

计算与语言 · 计算机科学 2023-08-17 Youxiang Zhu , Nana Lin , Xiaohui Liang , John A. Batsis , Robert M. Roth , Brian MacWhinney

In clinics, a radiology report is crucial for guiding a patient's treatment. However, writing radiology reports is a heavy burden for radiologists. To this end, we present an automatic, multi-modal approach for report generation from a…

图像与视频处理 · 电气工程与系统科学 2022-06-02 Shuxin Yang , Xian Wu , Shen Ge , S. Kevin Zhou , Li Xiao

Pneumonia has been one of the major causes of morbidities and mortality in the world and the prevalence of this disease is disproportionately high among the pediatric and elderly populations especially in resources trained areas Fast and…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Sathish Krishna Anumula , Vetrivelan Tamilmani , Aniruddha Arjun Singh , Dinesh Rajendran , Venkata Deepak Namburi

Searching for small objects in large images is a task that is both challenging for current deep learning systems and important in numerous real-world applications, such as remote sensing and medical imaging. Thorough scanning of very large…

计算机视觉与模式识别 · 计算机科学 2021-04-16 Nathan Drenkow , Philippe Burlina , Neil Fendley , Onyekachi Odoemene , Jared Markowitz

Recent research demonstrates that deep learning models are capable of precisely extracting bio-information (e.g. race, gender and age) from patients' Chest X-Rays (CXRs). In this paper, we further show that deep learning models are also…

图像与视频处理 · 电气工程与系统科学 2023-05-02 Hao Liang , Kevin Ni , Guha Balakrishnan

We propose a novel deep neural network architecture for normalcy detection in chest X-ray images. This architecture treats the problem as fine-grained binary classification in which the normal cases are well-defined as a class while leaving…

We propose a weakly-supervised approach that takes image-sentence pairs as input and learns to visually ground (i.e., localize) arbitrary linguistic phrases, in the form of spatial attention masks. Specifically, the model is trained with…

计算机视觉与模式识别 · 计算机科学 2017-05-04 Fanyi Xiao , Leonid Sigal , Yong Jae Lee

Bottom-up text detection methods play an important role in arbitrary-shape scene text detection but there are two restrictions preventing them from achieving their great potential, i.e., 1) the accumulation of false text segment detections,…

多媒体 · 计算机科学 2024-04-29 Chengpei Xu , Wenjing Jia , Ruomei Wang , Xiaonan Luo , Xiangjian He