English
Related papers

Related papers: Scaling Video Pretraining for Surgical Foundation …

200 papers

With the availability of large-scale, comprehensive, and general-purpose vision-language (VL) datasets such as MSCOCO, vision-language pre-training (VLP) has become an active area of research and proven to be effective for various VL tasks…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Li Xu , Bo Liu , Ameer Hamza Khan , Lu Fan , Xiao-Ming Wu

While self-supervised learning (SSL) algorithms have been widely used to pre-train deep models, few efforts [11] have been done to improve representation learning of X-ray image analysis with SSL pre-trained models. In this work, we study a…

Computer Vision and Pattern Recognition · Computer Science 2023-10-04 Weibin Liao , Haoyi Xiong , Qingzhong Wang , Yan Mo , Xuhong Li , Yi Liu , Zeyu Chen , Siyu Huang , Dejing Dou

Accurate segmentation and tracking of relevant elements of the surgical scene is crucial to enable context-aware intraoperative assistance and decision making. Current solutions remain tethered to domain-specific, supervised models that…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Jecia Z. Y. Mao , Francis X Creighton , Russell H Taylor , Manish Sahu

Video understanding has been considered as one critical step towards world modeling, which is an important long-term problem in AI research. Recently, multimodal foundation models have shown such potential via large-scale pretraining. These…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Boyu Chen , Siran Chen , Kunchang Li , Qinglin Xu , Yu Qiao , Yali Wang

This paper revisits the standard pretrain-then-finetune paradigm used in computer vision for visual recognition tasks. Typically, state-of-the-art foundation models are pretrained using large scale (weakly) supervised datasets with billions…

Training robust deep video representations has proven to be computationally challenging due to substantial decoding overheads, the enormous size of raw video streams, and their inherent high temporal redundancy. Different from existing…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Shristi Das Biswas , Efstathia Soufleri , Arani Roy , Kaushik Roy

The rapid growth of medical imaging has fueled the development of Foundation Models (FMs) to reduce the growing, unsustainable workload on radiologists. While recent FMs have shown the power of large-scale pre-training to CT and MRI…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Antoine Saporta , Baptiste Callard , Corentin Dancette , Julien Khlaut , Charles Corbière , Leo Butsanets , Amaury Prat , Pierre Manceron

Research in unpaired video translation has mainly focused on short-term temporal consistency by conditioning on neighboring frames. However for transfer from simulated to photorealistic sequences, available information on the underlying…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Dominik Rivoir , Micha Pfeiffer , Reuben Docea , Fiona Kolbinger , Carina Riediger , Jürgen Weitz , Stefanie Speidel

Foundation models leverage large-scale pretraining to capture extensive knowledge, demonstrating generalization in a wide range of language tasks. By comparison, vision foundation models (VFMs) often exhibit uneven improvements across…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Shiqi Huang , Yipei Wang , Natasha Thorley , Alexander Ng , Shaheer Saeed , Mark Emberton , Shonit Punwani , Veeru Kasivisvanathan , Dean Barratt , Daniel Alexander , Yipeng Hu

The remarkable zero-shot reasoning capabilities of large-scale Visual Language Models (VLMs) on static images have yet to be fully translated to the video domain. Conventional video understanding models often rely on extensive,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Shihao Ji , Zihui Song

Self-Supervised Learning (SSL) presents an exciting opportunity to unlock the potential of vast, untapped clinical datasets, for various downstream applications that suffer from the scarcity of labeled data. While SSL has revolutionized…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Tassilo Wald , Constantin Ulrich , Stanislav Lukyanenko , Andrei Goncharov , Alberto Paderno , Maximilian Miller , Leander Maerkisch , Paul F. Jäger , Klaus Maier-Hein

Despite recent successes, test-time scaling - i.e., dynamically expanding the token budget during inference as needed - remains brittle for vision-language models (VLMs): unstructured chains-of-thought about images entangle perception and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Niccolo Avogaro , Nayanika Debnath , Li Mi , Thomas Frick , Junling Wang , Zexue He , Hang Hua , Konrad Schindler , Mattia Rigotti

We introduce SurgFormer, a multiresolution gated transformer for data driven soft tissue simulation on volumetric meshes. High fidelity biomechanical solvers are often too costly for interactive use, so we train SurgFormer on solver…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Ashkan Shahbazi , Elaheh Akbari , Kyvia Pereira , Jon S. Heiselman , Annie C. Benson , Garrison L. H. Johnston , Jie Ying Wu , Nabil Simaan , Michael I. Miga , Soheil Kolouri

Automated endoscopy video analysis is a challenging task in medical computer vision, with the primary objective of assisting surgeons during procedures. The difficulty arises from the complexity of surgical scenes and the lack of a…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Dominik Batić , Felix Holm , Ege Özsoy , Tobias Czempiel , Nassir Navab

Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Tanzila Rahman , Renjie Liao , Leonid Sigal

In this work, we introduce SPLICE, a human-curated benchmark derived from the COIN instructional video dataset, designed to probe event-based reasoning across multiple dimensions: temporal, causal, spatial, contextual, and general…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Mohamad Ballout , Okajevo Wilfred , Seyedalireza Yaghoubi , Nohayr Muhammad Abdelmoneim , Julius Mayer , Elia Bruni

The scalability of current language-image pre-training for 3D medical imaging, such as CT and MRI, is constrained by the need for radiologists to manually curate raw clinical studies. In this work, we pioneer pre-training directly on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Chenhui Zhao , Yiwei Lyu , Asadur Chowdury , Edward Harake , Akhil Kondepudi , Akshay Rao , Xinhai Hou , Honglak Lee , Todd Hollon

Recorded videos from surgeries have become an increasingly important information source for the field of medical endoscopy, since the recorded footage shows every single detail of the surgery. However, while video recording is…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Sabrina Kletz , Klaus Schoeffmann , Jenny Benois-Pineau , Heinrich Husslein

Training a referring expression comprehension (ReC) model for a new visual domain requires collecting referring expressions, and potentially corresponding bounding boxes, for images in the domain. While large-scale pre-trained models are…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Sanjay Subramanian , William Merrill , Trevor Darrell , Matt Gardner , Sameer Singh , Anna Rohrbach

Large language models (LLMs) have shown strong reasoning capabilities when fine-tuned with reinforcement learning (RL). However, such methods require extensive data and compute, making them impractical under many realistic training budgets.…

Machine Learning · Computer Science 2026-04-17 Dai Do , Manh Nguyen , Svetha Venkatesh , Hung Le