English
Related papers

Related papers: From Phase Grounding to Intelligent Surgical Narra…

200 papers

Medical image segmentation remains challenging due to limited annotations for training, ambiguous anatomical features, and domain shifts. While vision-language models such as CLIP offer strong cross-modal representations, their potential…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Taha Koleilat , Hojat Asgariandehkordi , Omid Nejati Manzari , Berardino Barile , Yiming Xiao , Hassan Rivaz

Text-guided image generation aimed to generate desired images conditioned on given texts, while text-guided image manipulation refers to semantically edit parts of a given image based on specified texts. For these two similar tasks, the key…

Computer Vision and Pattern Recognition · Computer Science 2023-09-22 Xiaozhou You , Jian Zhang

Automatic recognition of surgical phases in surgical videos is a fundamental task in surgical workflow analysis. In this report, we propose a Transformer-based method that utilizes calibrated confidence scores for a 2-stage inference…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Yunfan Li , Vinayak Shenoy , Prateek Prasanna , I. V. Ramakrishnan , Haibin Ling , Himanshu Gupta

Human communication typically has an underlying structure. This is reflected in the fact that in many user generated videos, a starting point, ending, and certain objective steps between these two can be identified. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2016-01-28 Ozan Sener , Amir Zamir , Silvio Savarese , Ashutosh Saxena

Graph-based holistic scene representations facilitate surgical workflow understanding and have recently demonstrated significant success. However, this task is often hindered by the limited availability of densely annotated surgical scene…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Çağhan Köksal , Ghazal Ghazaei , Felix Holm , Azade Farshad , Nassir Navab

This paper proposes a practical multimodal video summarization task setting and a dataset to train and evaluate the task. The target task involves summarizing a given video into a predefined number of keyframe-caption pairs and displaying…

Computation and Language · Computer Science 2023-12-05 Keito Kudo , Haruki Nagasawa , Jun Suzuki , Nobuyuki Shimizu

Computer-assisted interventions can improve intra-operative guidance, particularly through deep learning methods that harness the spatiotemporal information in surgical videos. However, the severe data imbalance often found in surgical…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Danush Kumar Venkatesh , Isabel Funke , Micha Pfeiffer , Fiona Kolbinger , Hanna Maria Schmeiser , Juergen Weitz , Marius Distler , Stefanie Speidel

Camera-controlled video generation has made substantial progress, enabling generated videos to follow prescribed viewpoint trajectories. However, existing methods usually learn camera-specific conditioning through camera encoders, control…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yifan Wang , Tong He

Recent advancements in self-supervised learning have led to powerful surgical vision encoders capable of spatiotemporal understanding. However, extending these visual foundations to multi-modal reasoning tasks is severely bottlenecked by…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Chengan Che , Chao Wang , Jiayuan Huang , Xinyue Chen , Luis C. Garcia-Peraza-Herrera

Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes. Computer-assisted systems such as surgical visual question answering (VQA) offer promises for…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Shi Li , Vinkle Srivastav , Nicolas Chanel , Saurav Sharma , Nabani Banik , Lorenzo Arboit , Kun Yuan , Pietro Mascagni , Nicolas Padoy

Current video summarization methods rely heavily on supervised computer vision techniques, which demands time-consuming and subjective manual annotations. To overcome these limitations, we investigated self-supervised video summarization.…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Tomoya Sugihara , Shuntaro Masuda , Ling Xiao , Toshihiko Yamasaki

Large-scale vision-language models demonstrate strong multimodal alignment and generalization across diverse tasks. Among them, CLIP stands out as one of the most successful approaches. In this work, we extend the application of CLIP to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Sooyoung Park , Arda Senocak , Joon Son Chung

Surgical workflow recognition has numerous potential medical applications, such as the automatic indexing of surgical video databases and the optimization of real-time operating room scheduling, among others. As a result, phase recognition…

Computer Vision and Pattern Recognition · Computer Science 2016-05-24 Andru P. Twinanda , Sherif Shehata , Didier Mutter , Jacques Marescaux , Michel de Mathelin , Nicolas Padoy

Video recordings of open surgeries are greatly required for education and research purposes. However, capturing unobstructed videos is challenging since surgeons frequently block the camera field of view. To avoid occlusion, the positions…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Yuna Kato , Shohei Mori , Hideo Saito , Yoshifumi Takatsume , Hiroki Kajita , Mariko Isogawa

Understanding the workflow of surgical procedures in complex operating rooms requires a deep understanding of the interactions between clinicians and their environment. Surgical activity recognition (SAR) is a key computer vision task that…

Computer Vision and Pattern Recognition · Computer Science 2025-02-20 Idris Hamoud , Vinkle Srivastav , Muhammad Abdullah Jamal , Didier Mutter , Omid Mohareri , Nicolas Padoy

Video captioning is a critical task in the field of multimodal machine learning, aiming to generate descriptive and coherent textual narratives for video content. While large vision-language models (LVLMs) have shown significant progress,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Ji-jun Park , Soo-joon Choi

While foundation models have advanced surgical video analysis, current approaches rely predominantly on pixel-level reconstruction objectives that waste model capacity on low-level visual details, such as smoke, specular reflections, and…

Automated surgical workflow analysis and understanding can assist surgeons to standardize procedures and enhance post-surgical assessment and indexing, as well as, interventional monitoring. Computer-assisted interventional (CAI) systems…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Odysseas Zisimopoulos , Evangello Flouty , Imanol Luengo , Petros Giataganas , Jean Nehme , Andre Chow , Danail Stoyanov

VAR is a new generation paradigm that employs 'next-scale prediction' as opposed to 'next-token prediction'. This innovative transformation enables auto-regressive (AR) transformers to rapidly learn visual distributions and achieve robust…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Qian Zhang , Xiangzi Dai , Ninghua Yang , Xiang An , Ziyong Feng , Xingyu Ren

Video-Language Pre-training models have recently significantly improved various multi-modal downstream tasks. Previous dominant works mainly adopt contrastive learning to achieve global feature alignment across modalities. However, the…

Computer Vision and Pattern Recognition · Computer Science 2023-01-19 Fan Ma , Xiaojie Jin , Heng Wang , Jingjia Huang , Linchao Zhu , Jiashi Feng , Yi Yang