English
Related papers

Related papers: From Phase Grounding to Intelligent Surgical Narra…

200 papers

Human speech is often accompanied by body gestures including arm and hand gestures. We present a method that reenacts a high-quality video with gestures matching a target speech audio. The key idea of our method is to split and re-assemble…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Yang Zhou , Jimei Yang , Dingzeyu Li , Jun Saito , Deepali Aneja , Evangelos Kalogerakis

Surgical workflow anticipation can give predictions on what steps to conduct or what instruments to use next, which is an essential part of the computer-assisted intervention system for surgery, e.g. workflow reasoning in robotic surgery.…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Xiatian Zhang , Noura Al Moubayed , Hubert P. H. Shum

Physical computing infrastructure, data gathering, and algorithms have recently had significant advances to extract information from images and videos. The growth has been especially outstanding in image captioning and video captioning.…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Soheyla Amirian , Thiab R. Taha , Khaled Rasheed , Hamid R. Arabnia

Videos serve as a powerful medium to convey ideas, tell stories, and provide detailed instructions, especially through long-format tutorials. Such tutorials are valuable for learning new skills at one's own pace, yet they can be…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Nafisa Hussain

We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from most existing methods that only consider RGB images as…

Computer Vision and Pattern Recognition · Computer Science 2021-11-01 Yi-Wen Chen , Yi-Hsuan Tsai , Ming-Hsuan Yang

Surgical simulation plays a pivotal role in training novice surgeons, accelerating their learning curve and reducing intra-operative errors. However, conventional simulation tools fall short in providing the necessary photorealism and the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Ssharvien Kumar Sivakumar , Yannik Frisch , Ghazal Ghazaei , Anirban Mukhopadhyay

To develop intelligent speech assistants and integrate them seamlessly with intra-operative decision-support frameworks, accurate and efficient surgical phase recognition is a prerequisite. In this study, we propose a multimodal framework…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-24 Kubilay Can Demir , Belen Lojo Rodriguez , Tobias Weise , Andreas Maier , Seung Hee Yang

Video Question Answering (VideoQA) in the surgical domain aims to enhance intraoperative understanding by enabling AI models to reason over temporally coherent events rather than isolated frames. Current approaches are limited to static…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Mauro Orazio Drago , Luca Carlini , Pelinsu Celebi Balyemez , Dennis Pierantozzi , Chiara Lena , Cesare Hassan , Danail Stoyanov , Elena De Momi , Sophia Bano , Mobarak I. Hoque

Creating a vivid video from the event or scenario in our imagination is a truly fascinating experience. Recent advancements in text-to-video synthesis have unveiled the potential to achieve this with prompts only. While text is convenient…

Computer Vision and Pattern Recognition · Computer Science 2023-06-02 Jinbo Xing , Menghan Xia , Yuxin Liu , Yuechen Zhang , Yong Zhang , Yingqing He , Hanyuan Liu , Haoxin Chen , Xiaodong Cun , Xintao Wang , Ying Shan , Tien-Tsin Wong

Advances in surgical video analysis are transforming operating rooms into intelligent, data-driven environments. Computer-assisted systems support full surgical workflow, from preoperative planning to intraoperative guidance and…

Image and Video Processing · Electrical Eng. & Systems 2025-09-22 Sahar Nasirihaghighi

Modeling and recognition of surgical activities poses an interesting research problem. Although a number of recent works studied automatic recognition of surgical activities, generalizability of these works across different tasks and…

Computer Vision and Pattern Recognition · Computer Science 2020-08-17 Duygu Sarikaya , Pierre Jannin

The performance of deep learning (DL) algorithms is heavily influenced by the quantity and the quality of the annotated data. However, in Surgical Data Science, access to it is limited. It is thus unsurprising that substantial research…

Image and Video Processing · Electrical Eng. & Systems 2023-07-18 Luis C. Garcia-Peraza-Herrera , Sebastien Ourselin , Tom Vercauteren

Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Bohan Li , Shuojue Yang , Baorui Peng , Xianda Guo , Erli Zhang , Youqi Tao , Junfeng Duan , Daguang Xu , Qi Dou , Xin Jin , Wenjun Zeng , Hao Zhao , Yueming Jin

Vision-language models bridge visual and linguistic understanding and have proven to be powerful for video recognition tasks. Existing approaches primarily rely on parameter-efficient fine-tuning of image-text pre-trained models, yet they…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Wencheng Zhu , Yuexin Wang , Hongxuan Li , Pengfei Zhu , Qinghua Hu

Automatic instrument segmentation in video is an essentially fundamental yet challenging problem for robot-assisted minimally invasive surgery. In this paper, we propose a novel framework to leverage instrument motion information, by…

Computer Vision and Pattern Recognition · Computer Science 2019-07-19 Yueming Jin , Keyun Cheng , Qi Dou , Pheng-Ann Heng

Computer-assisted surgery (CAS) aims to provide the surgeon with the right type of assistance at the right moment. Such assistance systems are especially relevant in laparoscopic surgery, where CAS can alleviate some of the drawbacks that…

Computer Vision and Pattern Recognition · Computer Science 2017-02-14 Sebastian Bodenstedt , Martin Wagner , Darko Katić , Patrick Mietkowski , Benjamin Mayer , Hannes Kenngott , Beat Müller-Stich , Rüdiger Dillmann , Stefanie Speidel

Bridging vision and natural language is a longstanding goal in computer vision and multimedia research. While earlier works focus on generating a single-sentence description for visual content, recent works have studied paragraph…

Multimedia · Computer Science 2020-05-15 Junnan Li , Yongkang Wong , Qi Zhao , Mohan S. Kankanhalli

Multi-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury. To enable representation learning for medical images at scale, we turn to YouTube, a platform with a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Wisdom O. Ikezogwo , Kevin Zhang , Mehmet Saygin Seyfioglu , Fatemeh Ghezloo , Linda Shapiro , Ranjay Krishna

Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without relying on expert-labeled datasets. However, progress in…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Alejandra Perez , Chinedu Nwoye , Ramtin Raji Kermani , Omid Mohareri , Muhammad Abdullah Jamal

Long-form clinical videos are central to visual evidence-based decision-making, with growing importance for applications such as surgical robotics and related settings. However, current multimodal large language models typically process…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Wenjie Li , Yujie Zhang , Haoran Sun , Xingqi He , Hongcheng Gao , Chenglong Ma , Ming Hu , Guankun Wang , Shiyi Yao , Renhao Yang , Hongliang Ren , Lei Wang , Junjun He , Yankai Jiang