English
Related papers

Related papers: Scaling Video Pretraining for Surgical Foundation …

200 papers

Traditional open-access datasets focusing on surgical procedures are often limited by their small size, typically consisting of fewer than 100 videos and less than 30 hours of footage, which leads to poor model generalization. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Chengan Che , Chao Wang , Tom Vercauteren , Sophia Tsoka , Luis C. Garcia-Peraza-Herrera

Foundation models have achieved transformative success across biomedical domains by enabling holistic understanding of multimodal data. However, their application in surgery remains underexplored. Surgical intelligence presents unique…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Zhitao Zeng , Zhu Zhuo , Xiaojun Jia , Erli Zhang , Junde Wu , Jiaan Zhang , Yuxuan Wang , Chang Han Low , Jian Jiang , Zilong Zheng , Xiaochun Cao , Yutong Ban , Qi Dou , Yang Liu , Yueming Jin

Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without relying on expert-labeled datasets. However, progress in…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Alejandra Perez , Chinedu Nwoye , Ramtin Raji Kermani , Omid Mohareri , Muhammad Abdullah Jamal

Activity recognition in surgical videos is a key research area for developing next-generation devices and workflow monitoring systems. Since surgeries are long processes with highly-variable lengths, deep learning models used for surgical…

Computer Vision and Pattern Recognition · Computer Science 2022-09-08 Zhuohong He , Ali Mottaghi , Aidean Sharghi , Muhammad Abdullah Jamal , Omid Mohareri

Artificial intelligence, imaging, and large language models have the potential to transform surgical practice, training, and automation. Understanding and modeling of basic surgical actions (BSA), the fundamental unit of operation in any…

Capitalizing on image-level pre-trained models for various downstream tasks has recently emerged with promising performance. However, the paradigm of "image pre-training followed by video fine-tuning" for high-dimensional video data…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Shu Yang , Zhiyuan Cai , Luyang Luo , Ning Ma , Shuchang Xu , Hao Chen

Videos are prominent learning materials to prepare surgical trainees before they enter the operating room (OR). In this work, we explore techniques to enrich the video-based surgery learning experience. We propose Surgment, a system that…

Human-Computer Interaction · Computer Science 2024-06-27 Jingying Wang , Haoran Tang , Taylor Kantor , Tandis Soltani , Vitaliy Popov , Xu Wang

The automatic summarization of surgical videos is essential for enhancing procedural documentation, supporting surgical training, and facilitating post-operative analysis. This paper presents a novel method at the intersection of artificial…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Hugo Georgenthum , Cristian Cosentino , Fabrizio Marozzo , Pietro Liò

Surgical scene understanding is a cornerstone of computer-assisted intervention. While recent advances, particularly in surgical image segmentation, have driven progress, real-world clinical applications require a more holistic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Jincai Huang , Shihao Zou , Yuchen Guo , Jingjing Li , Wei Ji , Kai Wang , Shanshan Wang , Weixin Si

Surgical phase recognition is a critical component for context-aware decision support in intelligent operating rooms, yet training robust models is hindered by limited annotated clinical videos and large domain gaps between synthetic and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Yuxin He , An Li , Cheng Xue

The absence of openly accessible data and specialized foundation models is a major barrier for computational research in surgery. Toward this, (i) we open-source the largest dataset of general surgery videos to-date, consisting of 680 hours…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Samuel Schmidgall , Ji Woong Kim , Jeffrey Jopling , Axel Krieger

Multimodal large language models (LLMs) have achieved notable success across various domains, while research in the medical field has largely focused on unimodal images. Meanwhile, current general-domain multimodal models for videos still…

Computer Vision and Pattern Recognition · Computer Science 2024-08-16 Jiajie Li , Garrett Skinner , Gene Yang , Brian R Quaranto , Steven D Schwaitzberg , Peter C W Kim , Jinjun Xiong

Computer-assisted surgery research requires large, deeply annotated video datasets that capture clinical and technical variability. Existing cataract surgery resources lack the diversity and annotation depth required to train generalizable…

Surgical video-language pretraining (VLP) faces unique challenges due to the knowledge domain gap and the scarcity of multi-modal data. This study aims to bridge the gap by addressing issues regarding textual information loss in surgical…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Kun Yuan , Vinkle Srivastav , Nassir Navab , Nicolas Padoy

We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 billion curated image-text pairs and 10 billion video-hashtag…

We investigate how both the adaptation of a generic foundation model via transfer learning and the integration of complementary modalities from the operating room (OR) can support surgical data science. To this end, we use V-JEPA as the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Simon Pezold , Jérôme A. Kurylec , Jan S. Liechti , Beat P. Müller , Joël L. Lavanchy

Innovations in digital intelligence are transforming robotic surgery with more informed decision-making. Real-time awareness of surgical instrument presence and actions (e.g., cutting tissue) is essential for such systems. Yet, despite…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Jiajun Cheng , Xianwu Zhao , Sainan Liu , Xiaofan Yu , Ravi Prakash , Patrick J. Codd , Jonathan Elliott Katz , Shan Lin

Foundation models in video generation are demonstrating remarkable capabilities as potential world models for simulating the physical world. However, their application in high-stakes domains like surgery, which demand deep, specialized…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Zhen Chen , Qing Xu , Jinlin Wu , Biao Yang , Yuhao Zhai , Geng Guo , Jing Zhang , Yinlu Ding , Nassir Navab , Jiebo Luo

Geometric foundation models show promise in 3D reconstruction, yet their progress is severely constrained by the scarcity of diverse, large-scale 3D annotations. While Internet videos offer virtually unlimited raw data, utilizing them as a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Zihui Gao , Ke Liu , Donny Y. Chen , Duochao Shi , Guosheng Lin , Hao Chen , Chunhua Shen

Accurate robot segmentation is a fundamental capability for robotic perception. It enables precise visual servoing for VLA systems, scalable robot-centric data augmentation, accurate real-to-sim transfer, and reliable safety monitoring in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Haiyang Mei , Qiming Huang , Hai Ci , Mike Zheng Shou