English
Related papers

Related papers: Scaling Video Pretraining for Surgical Foundation …

200 papers

Computer-assisted surgery has been developed to enhance surgery correctness and safety. However, researchers and engineers suffer from limited annotated data to develop and train better algorithms. Consequently, the development of…

Computer Vision and Pattern Recognition · Computer Science 2020-12-24 W. -Y. Hong , C. -L. Kao , Y. -H. Kuo , J. -R. Wang , W. -L. Chang , C. -S. Shih

This paper presents an approach for surgical phase recognition using video data, aiming to provide a comprehensive understanding of surgical procedures for automated workflow analysis. The advent of robotic surgery, digitized operating…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Syed Abdul Mateen , Niharika Malvia , Syed Abdul Khader , Danny Wang , Deepti Srinivasan , Chi-Fu Jeffrey Yang , Lana Schumacher , Sandeep Manjanna

Reconstructing surgical scenes from monocular endoscopic video is critical for advancing robotic-assisted surgery. However, the application of state-of-the-art general-purpose reconstruction models is constrained by two key challenges: the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Kaiyuan Xu , Fangzhou Hong , Daniel Elson , Baoru Huang

We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus containing 36M high-quality video-caption pairs and 582M…

Learning meaningful visual representations in an embedding space can facilitate generalization in downstream tasks such as action segmentation and imitation. In this paper, we learn a motion-centric representation of surgical video…

Robotics · Computer Science 2020-06-02 Ajay Kumar Tanwani , Pierre Sermanet , Andy Yan , Raghav Anand , Mariano Phielipp , Ken Goldberg

Long-form video content constitutes a significant portion of internet traffic, making automated video summarization an essential research problem. However, existing video summarization datasets are notably limited in their size,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Dawit Mureja Argaw , Seunghyun Yoon , Fabian Caba Heilbron , Hanieh Deilamsalehy , Trung Bui , Zhaowen Wang , Franck Dernoncourt , Joon Son Chung

Video-language foundation models have proven to be highly effective in zero-shot applications across a wide range of tasks. A particularly challenging area is the intraoperative surgical procedure domain, where labeled data is scarce, and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Florian Stilz , Vinkle Srivastav , Nassir Navab , Nicolas Padoy

We introduce VidLPRO, a novel video-language (VL) pre-training framework designed specifically for robotic and laparoscopic surgery. While existing surgical VL models primarily rely on contrastive learning, we propose a more comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Mohammadmahdi Honarmand , Muhammad Abdullah Jamal , Omid Mohareri

Video-based assessments offer a scalable pathway for remote Parkinson's disease (PD) screening. While traditional approaches rely on handcrafted features mimicking clinical scales, recent advances in video foundation models (VFMs) enable…

Laparoscopic surgery is a complex surgical technique that requires extensive training. Recent advances in deep learning have shown promise in supporting this training by enabling automatic video-based assessment of surgical skills. However,…

In order to provide the right type of assistance at the right time, computer-assisted surgery systems need context awareness. To achieve this, methods for surgical workflow analysis are crucial. Currently, convolutional neural networks…

Computer Vision and Pattern Recognition · Computer Science 2018-10-05 Isabel Funke , Alexander Jenke , Sören Torge Mees , Jürgen Weitz , Stefanie Speidel , Sebastian Bodenstedt

Recent advances in multimodal large language models (LLMs) have highlighted their potential for medical and surgical applications. However, existing surgical datasets predominantly adopt a Visual Question Answering (VQA) format with…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Tae-Min Choi , Tae Kyeong Jeong , Garam Kim , Jaemin Lee , Yeongyoon Koh , In Cheul Choi , Jae-Ho Chung , Jong Woong Park , Juyoun Park

Surgical planning integrates visual perception, long-horizon reasoning, and procedural knowledge, yet it remains unclear whether current evaluation protocols reliably assess vision-language models (VLMs) in safety-critical settings.…

Computation and Language · Computer Science 2026-01-16 Ruochen Li , Kun Yuan , Yufei Xia , Yue Zhou , Qingyu Lu , Weihang Li , Youxiang Zhu , Nassir Navab

Multi-sequence Magnetic Resonance Imaging (MRI) offers remarkable versatility, enabling the distinct visualization of different tissue types. Nevertheless, the inherent heterogeneity among MRI sequences poses significant challenges to the…

Deep video recognition is more computationally expensive than image recognition, especially on large-scale datasets like Kinetics [1]. Therefore, training scalability is essential to handle a large amount of videos. In this paper, we study…

Computer Vision and Pattern Recognition · Computer Science 2019-12-10 Ji Lin , Chuang Gan , Song Han

Foundation models have exhibited remarkable success in various applications, such as disease diagnosis and text report generation. To date, a foundation model for endoscopic video analysis is still lacking. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Zhao Wang , Chang Liu , Shaoting Zhang , Qi Dou

Owing to recent advances in machine learning and the ability to harvest large amounts of data during robotic-assisted surgeries, surgical data science is ripe for foundational work. We present a large dataset of surgical videos and their…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Aneeq Zia , Max Berniker , Rogerio Nespolo , Xiaorui Zhang , Conor Perreault , Ziheng Wang , Benjamin Mueller , Ryan Schmidt , Kiran Bhattacharyya , Xi Liu , Anthony Jarc

Surgical Video Synthesis has emerged as a promising research direction following the success of diffusion models in general-domain video generation. Although existing approaches achieve high-quality video generation, most are unconditional…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Diego Biagini , Nassir Navab , Azade Farshad

Accurate surgical phase recognition is crucial for computer-assisted interventions and surgical video analysis. Annotating long surgical videos is labor-intensive, driving research toward leveraging unlabeled data for strong performance…

Scaling up model and data size have demonstrated impressive performance improvement over a wide range of tasks. Despite extensive studies on scaling behaviors for general-purpose tasks, medical images exhibit substantial differences from…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Jiarun Liu , Hong-Yu Zhou , Weijian Huang , Hao Yang , Dongning Song , Tao Tan , Yong Liang , Shanshan Wang