中文
相关论文

相关论文: GLeVE: Graph-Guided Lesion Grounding with Proposal…

200 篇论文

Building a large-scale training dataset is an essential problem in the development of medical image recognition systems. Visual grounding techniques, which automatically associate objects in images with corresponding descriptions, can…

计算机视觉与模式识别 · 计算机科学 2023-12-11 Akimichi Ichinose , Taro Hatsutani , Keigo Nakamura , Yoshiro Kitamura , Satoshi Iizuka , Edgar Simo-Serra , Shoji Kido , Noriyuki Tomiyama

Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign textual findings with visual evidence, leading to unreliable…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Pablo Messina , Andrés Villa , Juan León Alcázar , Karen Sánchez , Carlos Hinojosa , Denis Parra , Álvaro Soto , Bernard Ghanem

3D medical vision-language (VL) pretraining has shown potential in radiology by leveraging large-scale multimodal datasets with CT-report pairs. However, existing methods primarily rely on a global VL alignment directly adapted from 2D…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Jingyang Lin , Yingda Xia , Jianpeng Zhang , Ke Yan , Kai Cao , Le Lu , Jiebo Luo , Ling Zhang

Automated 3D segmentation of prostate lesions from biparametric MRI (bp-MRI) is essential for reliable algorithmic analysis, but achieving high precision remains challenging. Volumetric methods must combine multiple modalities while…

Vision-language models can connect the text description of an object to its specific location in an image through visual grounding. This has potential applications in enhanced radiology reporting. However, these models require large…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Zachary Huemann , Samuel Church , Joshua D. Warner , Daniel Tran , Xin Tie , Alan B McMillan , Junjie Hu , Steve Y. Cho , Meghan Lubner , Tyler J. Bradshaw

3D visual grounding aims to localize the unique target described by natural languages in 3D scenes. The significant gap between 3D and language modalities makes it a notable challenge to distinguish multiple similar objects through the…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Feng Xiao , Hongbin Xu , Guocan Zhao , Wenxiong Kang

Grounded video description (GVD) encourages captioning models to attend to appropriate video regions (e.g., objects) dynamically and generate a description. Such a setting can help explain the decisions of captioning models and prevents the…

计算机视觉与模式识别 · 计算机科学 2021-12-03 Wenqiao Zhang , Xin Eric Wang , Siliang Tang , Haizhou Shi , Haocheng Shi , Jun Xiao , Yueting Zhuang , William Yang Wang

3D ultrasound delivers high-resolution, real-time images of soft tissues, which is essential for pain research. However, manually distinguishing various tissues for quantitative analysis is labor-intensive. To streamline this process, we…

图像与视频处理 · 电气工程与系统科学 2025-03-26 Zixue Zeng , Xiaoyan Zhao , Matthew Cartier , Xin Meng , Jiantao Pu

Establishing dense correspondence across 3D shapes is crucial for fundamental downstream tasks, including texture transfer, shape interpolation, and robotic manipulation. However, learning these mappings without manual supervision remains a…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Qinfeng Xiao , Guofeng Mei , Qilong Liu , Chenyuan Yi , Fabio Poiesi , Jian Zhang , Bo Yang , Yick Kit-lun

Despite the progress of radiology report generation (RRG), existing works face two challenges: 1) The performances in clinical efficacy are unsatisfactory, especially for lesion attributes description; 2) the generated text lacks…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Haibo Jin , Haoxuan Che , Sunan He , Hao Chen

Radiology Report Generation (RRG) aims to produce accurate and coherent diagnostics from medical images. Although large vision language models (LVLM) improve report fluency and accuracy, they exhibit hallucinations, generating plausible yet…

计算与语言 · 计算机科学 2026-02-05 Ruixiao Yang , Yuanhe Tian , Xu Yang , Huiqi Li , Yan Song

Vision-language models (VLMs) have shown potential for automated radiology report generation, yet existing approaches rely on global embedding compression of volumetric data, often leading to hallucinated findings and limited anatomical…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Giuseppe A. Orlando , Paolo Papotti , Maria A. Zuluaga , Olivier Humbert , Marco Lorenzi

Current Vision-Language Models (VLMs) struggle to ground anatomical regions in 3D medical images and reason about them in a step-by-step manner, a key requirement of real-world diagnostic assessment. This ability is essential for aligning…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Sraavya Sambara , Sung Eun Kim , Xiaoman Zhang , Luyang Luo , Shreya Johri , Mohammed Baharoon , Du Hyun Ro , Pranav Rajpurkar

Vision--language models (VLMs) for radiology report generation (RRG) can produce long-form chest CT reports from volumetric scans and show strong potential to improve radiology workflow efficiency and consistency. However, existing methods…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Chenyu Wang , Weicheng Dai , Han Liu , Wenchao Li , Kayhan Batmanghelich

Clinically reliable perception of surgical scenes is essential for advancing intelligent, context-aware intraoperative assistance such as instrument handoff guidance, collision avoidance, and workflow-aware robotic support. Existing…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Tajamul Ashraf , Abrar Ul Riyaz , Wasif Tak , Tavaheed Tariq , Sonia Yadav , Moloud Abdar , Janibul Bashir

With the growing volume of CT examinations, there is an increasing demand for automated tools such as organ segmentation, abnormality detection, and report generation to support radiologists in managing their clinical workload. Multi-label…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Theo Di Piazza , Carole Lazarus , Olivier Nempont , Loic Boussel

Learning universal graph representations across heterogeneous domains is difficult because graph datasets differ in topology, node-attribute semantics, feature dimensions, and even attribute availability. We propose GraphVec, a…

机器学习 · 计算机科学 2026-05-08 Qi Feng , Jicong Fan

Visual-Semantic Embedding (VSE) is a prevalent approach in image-text retrieval by learning a joint embedding space between the image and language modalities where semantic similarities would be preserved. The triplet loss with…

计算机视觉与模式识别 · 计算机科学 2022-10-25 Hong Xuan , Xi Chen

Accurate lesion segmentation is essential in medical image analysis, yet most existing methods are designed for specific anatomical sites or imaging modalities, limiting their generalizability. Recent vision-language foundation models…

图像与视频处理 · 电气工程与系统科学 2026-03-30 Guoping Xu , Jayaram K. Udupa , Yubing Tong , Xin Long , Ying Zhang , Jie Deng , Weiguo Lu , You Zhang

We introduce a novel segmentation-aware joint training framework called generative reinforcement network (GRN) that integrates segmentation loss feedback to optimize both image generation and segmentation performance in a single stage. An…

‹ 上一页 1 2 3 10 下一页 ›