中文
相关论文

相关论文: A vision-language model and platform for temporall…

200 篇论文

Self-supervised representation learning has been highly promising for histopathology image analysis with numerous approaches leveraging their patient-slide-patch hierarchy to learn better representations. In this paper, we explore how the…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Hasindri Watawana , Kanchana Ranasinghe , Tariq Mahmood , Muzammal Naseer , Salman Khan , Fahad Shahbaz Khan

Recognizing various surgical tools, actions and phases from surgery videos is an important problem in computer vision with exciting clinical applications. Existing deep-learning-based methods for this problem either process each surgical…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Haifeng Wang , Hao Xu , Jun Wang , Jian Zhou , Ke Deng

We present VASTA, a novel vision and language-assisted Programming By Demonstration (PBD) system for smartphone task automation. Development of a robust PBD automation system requires overcoming three key challenges: first, how to make a…

Automated surgical workflow analysis is crucial for education, research, and clinical decision-making, but the lack of annotated datasets hinders the development of accurate and comprehensive workflow analysis solutions. We introduce a…

计算机视觉与模式识别 · 计算机科学 2025-03-17 David Gastager , Ghazal Ghazaei , Constantin Patsch

Medical Visual Question Answering (Med-VQA) holds significant potential for clinical decision support, yet existing efforts primarily focus on 2D imaging with limited task diversity. This paper presents 3D-RAD, a large-scale dataset…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Xiaotang Gai , Jiaxiang Liu , Yichen Li , Zijie Meng , Jian Wu , Zuozhu Liu

Modeling and automatically recognizing surgical activities are fundamental steps toward automation in surgery and play important roles in providing timely feedback to surgeons. Accurately recognizing surgical activities in video poses a…

图像与视频处理 · 电气工程与系统科学 2022-11-15 Abdishakour Awale , Duygu Sarikaya

We present RASO, a foundation model designed to Recognize Any Surgical Object, offering robust open-set recognition capabilities across a broad range of surgical procedures and object classes, in both surgical images and videos. RASO…

计算机视觉与模式识别 · 计算机科学 2025-05-07 Jiajie Li , Brian R Quaranto , Chenhui Xu , Ishan Mishra , Ruiyang Qin , Dancheng Liu , Peter C W Kim , Jinjun Xiong

While multi-modal foundation models pre-trained on large-scale data have been successful in natural language understanding and vision recognition, their use in medical domains is still limited due to the fine-grained nature of medical tasks…

计算机视觉与模式识别 · 计算机科学 2023-06-16 Xiaoman Zhang , Chaoyi Wu , Ya Zhang , Yanfeng Wang , Weidi Xie

In our pursuit of advancing multi-modal AI assistants capable of guiding users to achieve complex multi-step goals, we propose the task of "Visual Planning for Assistance (VPA)". Given a succinct natural language goal, e.g., "make a shelf",…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Dhruvesh Patel , Hamid Eghbalzadeh , Nitin Kamra , Michael Louis Iuzzolino , Unnat Jain , Ruta Desai

To meet the growing demand for systematic surgical training, wet-lab environments have become indispensable platforms for hands-on practice in ophthalmology. Yet, traditional wet-lab training depends heavily on manual performance…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Negin Ghamsarian , Raphael Sznitman , Klaus Schoeffmann , Jens Kowal

This paper presents a novel approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot…

Laparoscopic surgery constrains surgeons spatial awareness because procedures are performed through a monocular, two-dimensional (2D) endoscopic view. Conventional training methods using dry-lab models or recorded videos provide limited…

人机交互 · 计算机科学 2025-11-05 Songyang Liu , Yunpeng Tan , Shuai Li

Surgical video understanding is a crucial prerequisite for advancing Computer-Assisted Surgery. While vision-language models (VLMs) have recently been applied to the surgical domain, existing surgical vision-language datasets lack in…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Lennart Maack , Alexander Schlaefer

We introduce Being-H0, a dexterous Vision-Language-Action model (VLA) trained on large-scale human videos. Existing VLAs struggle with complex manipulation tasks requiring high dexterity and generalize poorly to novel scenarios and tasks,…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Hao Luo , Yicheng Feng , Wanpeng Zhang , Sipeng Zheng , Ye Wang , Haoqi Yuan , Jiazheng Liu , Chaoyi Xu , Qin Jin , Zongqing Lu

Real-time tool segmentation from endoscopic videos is an essential part of many computer-assisted robotic surgical systems and of critical importance in robotic surgical data science. We propose two novel deep learning architectures for…

Spatiotemporal reasoning is a fundamental capability for artificial intelligence (AI) in soft tissue surgery, paving the way for intelligent assistive systems and autonomous robotics. While 2D vision-language models show increasing promise…

Video understanding of robot-assisted surgery (RAS) videos is an active research area. Modeling the gestures and skill level of surgeons presents an interesting problem. The insights drawn may be applied in effective skill acquisition,…

计算机视觉与模式识别 · 计算机科学 2020-08-04 Duygu Sarikaya , Jason J. Corso , Khurshid A. Guru

Every day, countless surgeries are performed worldwide, each within the distinct settings of operating rooms (ORs) that vary not only in their setups but also in the personnel, tools, and equipment used. This inherent diversity poses a…

计算机视觉与模式识别 · 计算机科学 2024-04-11 Ege Özsoy , Chantal Pellegrini , Matthias Keicher , Nassir Navab

Open, or non-laparoscopic surgery, represents the vast majority of all operating room procedures, but few tools exist to objectively evaluate these techniques at scale. Current efforts involve human expert-based visual assessment. We…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Michael Zhang , Xiaotian Cheng , Daniel Copeland , Arjun Desai , Melody Y. Guan , Gabriel A. Brat , Serena Yeung

3D Gaussian Splatting offers expressive scene reconstruction, modeling a broad range of visual, geometric, and semantic information. However, efficient real-time map reconstruction with data streamed from multiple robots and devices remains…

机器人学 · 计算机科学 2025-06-04 Javier Yu , Timothy Chen , Mac Schwager