English
Related papers

Related papers: GroundedSurg: A Multi-Procedure Benchmark for Lang…

200 papers

Surgery is a highly complex process, and artificial intelligence has emerged as a transformative force in supporting surgical guidance and decision-making. However, the unimodal nature of most current AI systems limits their ability to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Nakul Poudel , Richard Simon , Cristian A. Linte

Grounding radiology report descriptions to 3D CT volumes is essential for verifiable clinical interpretation, yet remains challenging due to the semantic-spatial gap between free-text narratives and volumetric anatomy. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Shuo Jiang , Yuhao Hong , Chunbo Jiang , Weihong Chen , Huangwei Chen , Shenghao Zhu , Beining Wu , Mingxuan Liu , Zhu Zhu , Feiwei Qin , Min Tan , Yifei Chen

Video reasoning segmentation requires localizing objects across video frames from natural language expressions, often involving spatial reasoning and implicit references. Recent approaches leverage frozen large vision-language models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Ali Cheraghian , Hamidreza Dastmalchi , Abdelwahed Khamis , Morteza Saberi , Aijun An , Lars Petersson

Real-time tool segmentation from endoscopic videos is an essential part of many computer-assisted robotic surgical systems and of critical importance in robotic surgical data science. We propose two novel deep learning architectures for…

Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression text. This task simulates the common referential relationships…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Linhui Xiao , Xiaoshan Yang , Xiangyuan Lan , Yaowei Wang , Changsheng Xu

Enabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizing surgical…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Meng Wei , Kun Yuan , Shi Li , Yue Zhou , Long Bai , Nassir Navab , Hongliang Ren , Hong Joo Lee , Tom Vercauteren , Nicolas Padoy

Existing visual grounding benchmarks primarily evaluate alignment between image regions and literal referring expressions, where models can often succeed by matching a prominent named category. We explore a complementary and more…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Ruozhen He , Nisarg A. Shah , Qihua Dong , Zilin Xiao , Jaywon Koo , Vicente Ordonez

Referring expression grounding is a core problem in visual grounding and is widely used as a diagnostic of spatial grounding and reasoning in vision and language models, yet most prior work focuses on natural images. In contrast, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Tianhao Niu , Ziyu Han , Qingfu Zhu , Wanxiang Che

Surgical scene understanding is critical for surgical training and robotic decision-making in robot-assisted surgery. Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated great potential for advancing scene…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Guankun Wang , Junyi Wang , Wenjin Mo , Long Bai , Kun Yuan , Ming Hu , Jinlin Wu , Junjun He , Yiming Huang , Nicolas Padoy , Zhen Lei , Hongbin Liu , Nassir Navab , Hongliang Ren

The detection and grounding of manipulated content in multimodal data has emerged as a critical challenge in media forensics. While existing benchmarks demonstrate technical progress, they suffer from misalignment artifacts that poorly…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Jinjie Shen , Yaxiong Wang , Lechao Cheng , Nan Pu , Zhun Zhong

Chronic wounds affect a large population, particularly the elderly and diabetic patients, who often exhibit limited mobility and co-existing health conditions. Automated wound monitoring via mobile image capture can reduce in-person…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Vanessa Borst , Timo Dittus , Tassilo Dege , Astrid Schmieder , Samuel Kounev

Endoscopic procedures are essential for diagnosing and treating internal diseases, and multi-modal large language models (MLLMs) are increasingly applied to assist in endoscopy analysis. However, current benchmarks are limited, as they…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Shengyuan Liu , Boyun Zheng , Wenting Chen , Zhihao Peng , Zhenfei Yin , Jing Shao , Jiancong Hu , Yixuan Yuan

Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric benchmark targeting low-level spatial intelligence in embodied 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Jiyao Zhang , Mingxu Zhang , Yitong Peng , Haoxuan Liu , Chenshuo Wang , Yuxing Long , Haoyang Huang , Dongjiang Li , Nan Duan , Hui Shen , Hao Dong

Robot-assisted neurological surgery is receiving growing interest due to the improved dexterity, precision, and control of surgical tools, which results in better patient outcomes. However, such systems often limit surgeons' natural sensory…

Signal Processing · Electrical Eng. & Systems 2025-08-13 Zacharias Chen , Alexa Cristelle Cahilig , Sarah Dias , Prithu Kolar , Ravi Prakash , Patrick J. Codd

Most multimodal large language models (MLLMs) learn language-to-object grounding through causal language modeling where grounded objects are captured by bounding boxes as sequences of location tokens. This paradigm lacks pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Yichi Zhang , Ziqiao Ma , Xiaofeng Gao , Suhaila Shakiah , Qiaozi Gao , Joyce Chai

We introduce Grounded Situation Recognition (GSR), a task that requires producing structured semantic summaries of images describing: the primary activity, entities engaged in the activity with their roles (e.g. agent, tool), and…

Computer Vision and Pattern Recognition · Computer Science 2020-03-27 Sarah Pratt , Mark Yatskar , Luca Weihs , Ali Farhadi , Aniruddha Kembhavi

Object rearrangement has recently emerged as a key competency in robot manipulation, with practical solutions generally involving object detection, recognition, grasping and high-level planning. Goal-images describing a desired scene…

Robotics · Computer Science 2021-11-16 Walter Goodwin , Sagar Vaze , Ioannis Havoutis , Ingmar Posner

Accurate medical image segmentation is essential for clinical diagnosis and treatment planning. While recent interactive foundation models (e.g., nnInteractive) enhance generalization through large-scale multimodal pretraining, they still…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Ziyu Zhang , Yi Yu , Simeng Zhu , Ahmed Aly , Yunhe Gao , Ning Gu , Yuan Xue

Most state-of-the-art techniques for medical image segmentation rely on deep-learning models. These models, however, are often trained on narrowly-defined tasks in a supervised fashion, which requires expensive labeled datasets. Recent…

Image and Video Processing · Electrical Eng. & Systems 2023-10-04 Heejong Kim , Victor Ion Butoi , Adrian V. Dalca , Daniel J. A. Margolis , Mert R. Sabuncu

Visual grounding, localizing objects from natural language descriptions, represents a critical bridge between language and vision understanding. While multimodal large language models (MLLMs) achieve impressive scores on existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Rang Li , Lei Li , Shuhuai Ren , Hao Tian , Shuhao Gu , Shicheng Li , Zihao Yue , Yudong Wang , Wenhan Ma , Zhe Yang , Jingyuan Ma , Zhifang Sui , Fuli Luo
‹ Prev 1 3 4 5 6 7 10 Next ›