English
Related papers

Related papers: A vision-language model and platform for temporall…

200 papers

Self-supervised representation learning has been highly promising for histopathology image analysis with numerous approaches leveraging their patient-slide-patch hierarchy to learn better representations. In this paper, we explore how the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Hasindri Watawana , Kanchana Ranasinghe , Tariq Mahmood , Muzammal Naseer , Salman Khan , Fahad Shahbaz Khan

Recognizing various surgical tools, actions and phases from surgery videos is an important problem in computer vision with exciting clinical applications. Existing deep-learning-based methods for this problem either process each surgical…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Haifeng Wang , Hao Xu , Jun Wang , Jian Zhou , Ke Deng

We present VASTA, a novel vision and language-assisted Programming By Demonstration (PBD) system for smartphone task automation. Development of a robust PBD automation system requires overcoming three key challenges: first, how to make a…

Human-Computer Interaction · Computer Science 2019-11-06 Alborz Rezazadeh Sereshkeh , Gary Leung , Krish Perumal , Caleb Phillips , Minfan Zhang , Afsaneh Fazly , Iqbal Mohomed

Automated surgical workflow analysis is crucial for education, research, and clinical decision-making, but the lack of annotated datasets hinders the development of accurate and comprehensive workflow analysis solutions. We introduce a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 David Gastager , Ghazal Ghazaei , Constantin Patsch

Medical Visual Question Answering (Med-VQA) holds significant potential for clinical decision support, yet existing efforts primarily focus on 2D imaging with limited task diversity. This paper presents 3D-RAD, a large-scale dataset…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Xiaotang Gai , Jiaxiang Liu , Yichen Li , Zijie Meng , Jian Wu , Zuozhu Liu

Modeling and automatically recognizing surgical activities are fundamental steps toward automation in surgery and play important roles in providing timely feedback to surgeons. Accurately recognizing surgical activities in video poses a…

Image and Video Processing · Electrical Eng. & Systems 2022-11-15 Abdishakour Awale , Duygu Sarikaya

We present RASO, a foundation model designed to Recognize Any Surgical Object, offering robust open-set recognition capabilities across a broad range of surgical procedures and object classes, in both surgical images and videos. RASO…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Jiajie Li , Brian R Quaranto , Chenhui Xu , Ishan Mishra , Ruiyang Qin , Dancheng Liu , Peter C W Kim , Jinjun Xiong

While multi-modal foundation models pre-trained on large-scale data have been successful in natural language understanding and vision recognition, their use in medical domains is still limited due to the fine-grained nature of medical tasks…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Xiaoman Zhang , Chaoyi Wu , Ya Zhang , Yanfeng Wang , Weidi Xie

In our pursuit of advancing multi-modal AI assistants capable of guiding users to achieve complex multi-step goals, we propose the task of "Visual Planning for Assistance (VPA)". Given a succinct natural language goal, e.g., "make a shelf",…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Dhruvesh Patel , Hamid Eghbalzadeh , Nitin Kamra , Michael Louis Iuzzolino , Unnat Jain , Ruta Desai

To meet the growing demand for systematic surgical training, wet-lab environments have become indispensable platforms for hands-on practice in ophthalmology. Yet, traditional wet-lab training depends heavily on manual performance…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Negin Ghamsarian , Raphael Sznitman , Klaus Schoeffmann , Jens Kowal

This paper presents a novel approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot…

Laparoscopic surgery constrains surgeons spatial awareness because procedures are performed through a monocular, two-dimensional (2D) endoscopic view. Conventional training methods using dry-lab models or recorded videos provide limited…

Human-Computer Interaction · Computer Science 2025-11-05 Songyang Liu , Yunpeng Tan , Shuai Li

Surgical video understanding is a crucial prerequisite for advancing Computer-Assisted Surgery. While vision-language models (VLMs) have recently been applied to the surgical domain, existing surgical vision-language datasets lack in…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Lennart Maack , Alexander Schlaefer

We introduce Being-H0, a dexterous Vision-Language-Action model (VLA) trained on large-scale human videos. Existing VLAs struggle with complex manipulation tasks requiring high dexterity and generalize poorly to novel scenarios and tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Hao Luo , Yicheng Feng , Wanpeng Zhang , Sipeng Zheng , Ye Wang , Haoqi Yuan , Jiazheng Liu , Chaoyi Xu , Qin Jin , Zongqing Lu

Real-time tool segmentation from endoscopic videos is an essential part of many computer-assisted robotic surgical systems and of critical importance in robotic surgical data science. We propose two novel deep learning architectures for…

Spatiotemporal reasoning is a fundamental capability for artificial intelligence (AI) in soft tissue surgery, paving the way for intelligent assistive systems and autonomous robotics. While 2D vision-language models show increasing promise…

Video understanding of robot-assisted surgery (RAS) videos is an active research area. Modeling the gestures and skill level of surgeons presents an interesting problem. The insights drawn may be applied in effective skill acquisition,…

Computer Vision and Pattern Recognition · Computer Science 2020-08-04 Duygu Sarikaya , Jason J. Corso , Khurshid A. Guru

Every day, countless surgeries are performed worldwide, each within the distinct settings of operating rooms (ORs) that vary not only in their setups but also in the personnel, tools, and equipment used. This inherent diversity poses a…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Ege Özsoy , Chantal Pellegrini , Matthias Keicher , Nassir Navab

Open, or non-laparoscopic surgery, represents the vast majority of all operating room procedures, but few tools exist to objectively evaluate these techniques at scale. Current efforts involve human expert-based visual assessment. We…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Michael Zhang , Xiaotian Cheng , Daniel Copeland , Arjun Desai , Melody Y. Guan , Gabriel A. Brat , Serena Yeung

3D Gaussian Splatting offers expressive scene reconstruction, modeling a broad range of visual, geometric, and semantic information. However, efficient real-time map reconstruction with data streamed from multiple robots and devices remains…

Robotics · Computer Science 2025-06-04 Javier Yu , Timothy Chen , Mac Schwager