English
Related papers

Related papers: Anatomical grounding pre-training for medical phra…

200 papers

Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos. In medical imaging research, VLMs represent a bridge between object detection and segmentation, and report understanding and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Andrew Seohwan Yu , Mohsen Hariri , Kunio Nakamura , Mingrui Yang , Xiaojuan Li , Vipin Chaudhary

Key to tasks that require reasoning about natural language in visual contexts is grounding words and phrases to image regions. However, observing this grounding in contemporary models is complex, even if it is generally expected to take…

Computation and Language · Computer Science 2024-06-03 Noriyuki Kojima , Hadar Averbuch-Elor , Yoav Artzi

Vision-language models (VLMs) have recently shown remarkable zero-shot performance in medical image understanding, yet their grounding ability, the extent to which textual concepts align with visual evidence, remains underexplored. In the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Haozhe Luo , Shelley Zixin Shu , Ziyu Zhou , Sebastian Otalora , Mauricio Reyes

Medical imaging is widely used in clinical practice for diagnosis and treatment. Report-writing can be error-prone for unexperienced physicians, and time- consuming and tedious for experienced physicians. To address these issues, we study…

Computation and Language · Computer Science 2019-01-09 Baoyu Jing , Pengtao Xie , Eric Xing

Deformable image registration is a fundamental problem in the field of medical image analysis. During the last years, we have witnessed the advent of deep learning-based image registration methods which achieve state-of-the-art performance,…

Image and Video Processing · Electrical Eng. & Systems 2020-02-03 Lucas Mansilla , Diego H. Milone , Enzo Ferrante

Visual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Ta Duc Huy , Duy Anh Huynh , Yutong Xie , Yuankai Qi , Qi Chen , Phi Le Nguyen , Sen Kim Tran , Son Lam Phung , Anton van den Hengel , Zhibin Liao , Minh-Son To , Johan W. Verjans , Vu Minh Hieu Phan

Cardiac ultrasound diagnosis is critical for cardiovascular disease assessment, but acquiring standard views remains highly operator-dependent. Existing medical segmentation models often yield anatomically inconsistent results in images…

Robotics · Computer Science 2026-03-24 Zhiyan Cao , Zhengxi Wu , Yiwei Wang , Pei-Hsuan Lin , Li Zhang , Zhen Xie , Huan Zhao , Han Ding

Neuroscientists have recently turned to intracranial brain recording methods, like electrocorticography (ECoG), for human experiments because of the fine spatial and temporal resolution that they afford. Models trained on this data,…

Computation and Language · Computer Science 2026-05-20 Aditya R. Vaidya , Richard J. Antonello , Alexander G. Huth

Object segmentation plays an important role in the modern medical image analysis, which benefits clinical study, disease diagnosis, and surgery planning. Given the various modalities of medical images, the automated or semi-automated…

Image and Video Processing · Electrical Eng. & Systems 2020-06-01 Dong Yang , Holger Roth , Xiaosong Wang , Ziyue Xu , Andriy Myronenko , Daguang Xu

Medical image segmentation, the task of partitioning an image into meaningful parts, is an important step toward automating medical image analysis and is at the crux of a variety of medical imaging applications, such as computer aided…

Computer Vision and Pattern Recognition · Computer Science 2016-07-06 Masoud S. Nosrati , Ghassan Hamarneh

Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit major competency gaps in multimodal understanding and reasoning especially in high-value verticals…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Qianchu Liu , Sheng Zhang , Guanghui Qin , Yu Gu , Ying Jin , Sam Preston , Yanbo Xu , Sid Kiblawi , Wen-wai Yim , Tim Ossowski , Tristan Naumann , Mu Wei , Hoifung Poon

Pathological structures in medical images are typically deviations from the expected anatomy of a patient. While clinicians consider this interplay between anatomy and pathology, recent deep learning algorithms specialize in recognizing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Alexander Jaus , Constantin Seibold , Simon Reiß , Lukas Heine , Anton Schily , Moon Kim , Fin Hendrik Bahnsen , Ken Herrmann , Rainer Stiefelhagen , Jens Kleesiek

In this paper, we introduce a contextual grounding approach that captures the context in corresponding text entities and image regions to improve the grounding accuracy. Specifically, the proposed architecture accepts pre-trained text token…

Computer Vision and Pattern Recognition · Computer Science 2019-11-07 Farley Lai , Ning Xie , Derek Doran , Asim Kadav

Temporal Video Grounding (TVG) aims to localize a moment from an untrimmed video given the language description. Since the annotation of TVG is labor-intensive, TVG under limited supervision has accepted attention in recent years. The great…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Xing Zhang , Jiaxi Gu , Haoyu Zhao , Shicong Wang , Hang Xu , Renjing Pei , Songcen Xu , Zuxuan Wu , Yu-Gang Jiang

Temporal comparison of chest X-rays is fundamental to clinical radiology, enabling detection of disease progression, treatment response, and new findings. While vision-language models have advanced single-image report generation and visual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 OFM Riaz Rahman Aranya , Kevin Desai

We present ASAP, a new framework for detecting and grounding multi-modal media manipulation (DGM4).Upon thorough examination, we observe that accurate fine-grained cross-modal semantic alignment between the image and text is vital for…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Zhenxing Zhang , Yaxiong Wang , Lechao Cheng , Zhun Zhong , Dan Guo , Meng Wang

The construction of 3D medical image datasets presents several issues, including requiring significant financial costs in data collection and specialized expertise for annotation, as well as strict privacy concerns for patient…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Ryu Tadokoro , Ryosuke Yamada , Kodai Nakashima , Ryo Nakamura , Hirokatsu Kataoka

Endowing Large Multimodal Models (LMMs) with visual grounding capability can significantly enhance AIs' understanding of the visual world and their interaction with humans. However, existing methods typically fine-tune the parameters of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Size Wu , Sheng Jin , Wenwei Zhang , Lumin Xu , Wentao Liu , Wei Li , Chen Change Loy

Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging implicit geometric priors from 3D reconstruction models,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jiaxin Zhang , Junjun Jiang , Haijie Li , Youyu Chen , Kui Jiang , Dave Zhenyu Chen

Large-scale pre-training holds the promise to advance 3D medical object detection, a crucial component of accurate computer-aided diagnosis. Yet, it remains underexplored compared to segmentation, where pre-training has already demonstrated…

Image and Video Processing · Electrical Eng. & Systems 2025-09-22 Katharina Eckstein , Constantin Ulrich , Michael Baumgartner , Jessica Kächele , Dimitrios Bounias , Tassilo Wald , Ralf Floca , Klaus H. Maier-Hein