中文
相关论文

相关论文: GLeVE: Graph-Guided Lesion Grounding with Proposal…

200 篇论文

Due to imaging artifacts and low signal-to-noise ratio in ultrasound images, automatic bone surface segmentation networks often produce fragmented predictions that can hinder the success of ultrasound-guided computer-assisted surgical…

图像与视频处理 · 电气工程与系统科学 2022-06-20 Aimon Rahman , Wele Gedara Chaminda Bandara , Jeya Maria Jose Valanarasu , Ilker Hacihaliloglu , Vishal M Patel

Medical image grounding aims to align natural language phrases with specific regions in medical images, serving as a foundational task for intelligent diagnosis, visual question answering (VQA), and automated report generation (MRG).…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Ziye Deng , Ruihan He , Jiaxiang Liu , Yuan Wang , Zijie Meng , Songtao Jiang , Yong Xie , Zuozhu Liu

Embedding methods transform the knowledge graph into a continuous, low-dimensional space, facilitating inference and completion tasks. Existing methods are mainly divided into two types: translational distance models and semantic matching…

信息检索 · 计算机科学 2025-03-11 Deepak Banerjee , Anjali Ishaan

3D visual grounding aims to automatically locate the 3D region of the specified object given the corresponding textual description. Existing works fail to distinguish similar objects especially when multiple referred objects are involved in…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Feng Xiao , Hongbin Xu , Qiuxia Wu , Wenxiong Kang

3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a free-form language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and…

计算机视觉与模式识别 · 计算机科学 2021-03-31 Mingtao Feng , Zhen Li , Qi Li , Liang Zhang , XiangDong Zhang , Guangming Zhu , Hui Zhang , Yaonan Wang , Ajmal Mian

Grounding referring expressions in RGBD image has been an emerging field. We present a novel task of 3D visual grounding in single-view RGBD image where the referred objects are often only partially scanned due to occlusion. In contrast to…

计算机视觉与模式识别 · 计算机科学 2021-03-18 Haolin Liu , Anran Lin , Xiaoguang Han , Lei Yang , Yizhou Yu , Shuguang Cui

Objective: To evaluate large language models (LLMs) against supervised baselines for fine-grained, lesion-level detection of incidentalomas requiring follow-up, addressing the limitations of current document-level classification systems.…

计算与语言 · 计算机科学 2025-12-08 Namu Park , Farzad Ahmed , Zhaoyi Sun , Kevin Lybarger , Ethan Breinhorst , Julie Hu , Ozlem Uzuner , Martin Gunn , Meliha Yetisgen

Developing generalist foundation model has recently attracted tremendous attention among researchers in the field of AI for Medicine (AI4Medicine). A pivotal insight in developing these models is their reliance on dataset scaling, which…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Xiaoman Zhang , Chaoyi Wu , Ziheng Zhao , Jiayu Lei , Ya Zhang , Yanfeng Wang , Weidi Xie

Vision-language models (VLMs) read an image and produce text in a single forward pass, whereas radiologists typically inspect an image several times and consult the literature before writing a report. We introduce GAZE (Grounded Agentic…

机器学习 · 计算机科学 2026-05-05 Duaa Alim , Mogtaba Alim , Liam Chalcroft

Vision-Language Models (VLMs) can generate convincing clinical narratives, yet frequently struggle to visually ground their statements. We posit this limitation arises from the scarcity of high-quality, large-scale clinical…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Mengmeng Zhang , Xiaoping Wu , Hao Luo , Fan Wang , Yisheng Lv

Radiology reporting generative AI holds significant potential to alleviate clinical workloads and streamline medical care. However, achieving high clinical accuracy is challenging, as radiological images often feature subtle lesions and…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Yijian Gao , Dominic Marshall , Xiaodan Xing , Junzhi Ning , Giorgos Papanastasiou , Guang Yang , Matthieu Komorowski

Parse graphs boost human pose estimation (HPE) by integrating context and hierarchies, yet prior work mostly focuses on single modality modeling, ignoring the potential of multimodal fusion. Notably, language offers rich HPE priors like…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Shibang Liu , Xuemei Xie , Guangming Shi

The rapid growth of AI-generated imagery has blurred the boundary between real and synthetic content, raising practical concerns for digital integrity. Vision-language models (VLMs) can provide natural language explanations, but standard…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Yikun Ji , Yan Hong , Bowen Deng , Jun Lan , Huijia Zhu , Weiqiang Wang , Liqing Zhang , Jianfu Zhang

Self-supervised representation learning on text-attributed graphs, which aims to create expressive and generalizable representations for various downstream tasks, has received increasing research attention lately. However, existing methods…

计算与语言 · 计算机科学 2023-10-24 Yichuan Li , Kaize Ding , Kyumin Lee

Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VLM) approaches achieve promising results but rely on view-dependent reasoning or implicit…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Xuefei Sun , Xujia Zhang , Brendan Crowe , Doncey Albin , Christoffer Heckman

Temporal comparison of chest X-rays is fundamental to clinical radiology, enabling detection of disease progression, treatment response, and new findings. While vision-language models have advanced single-image report generation and visual…

计算机视觉与模式识别 · 计算机科学 2026-02-04 OFM Riaz Rahman Aranya , Kevin Desai

We present a structural graph reasoning framework that incorporates explicit anatomical priors for explainable vision-based diagnosis. Convolutional feature maps are reinterpreted as patch-level graphs, where nodes encode both appearance…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Khaled Berkani

Models like OpenAI-o3 pioneer visual grounded reasoning by dynamically referencing visual regions, just like human "thinking with images". However, no benchmark exists to evaluate these capabilities holistically. To bridge this gap, we…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Haochen Wang , Xiangtai Li , Zilong Huang , Anran Wang , Jiacong Wang , Tao Zhang , Jiani Zheng , Sule Bai , Zijian Kang , Jiashi Feng , Zhuochen Wang , Zhaoxiang Zhang

In radiologists' routine work, one major task is to read a medical image, e.g., a CT scan, find significant lesions, and write sentences in the radiology report to describe them. In this paper, we study the lesion description or annotation…

计算机视觉与模式识别 · 计算机科学 2019-03-28 Ke Yan , Yifan Peng , Zhiyong Lu , Ronald M. Summers

Radiology Report Generation (RRG) aims to automatically generate diagnostic reports from radiology images. To achieve this, existing methods have leveraged the powerful cross-modal generation capabilities of Multimodal Large Language Models…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Jiechao Gao , Chang Liu , Yuangang Li