English
Related papers

Related papers: SurgOnAir: Hierarchy-Aware Real-Time Surgical Vide…

200 papers

Surgical workflow anticipation is the task of predicting the timing of relevant surgical events from live video data, which is critical in Robotic-Assisted Surgery (RAS). Accurate predictions require the use of spatial information to model…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Francis Xiatian Zhang , Jingjing Deng , Robert Lieck , Hubert P. H. Shum

Laparoscopic surgery constrains surgeons spatial awareness because procedures are performed through a monocular, two-dimensional (2D) endoscopic view. Conventional training methods using dry-lab models or recorded videos provide limited…

Human-Computer Interaction · Computer Science 2025-11-05 Songyang Liu , Yunpeng Tan , Shuai Li

Surgical intelligence has the potential to improve the safety and consistency of surgical care, yet most existing surgical AI frameworks remain task-specific and struggle to generalize across procedures and institutions. Although multimodal…

We introduce a speech-guided embodied agent framework for video-guided skull base surgery that dynamically executes perception and image-guidance tasks in response to surgeon queries. The proposed system integrates natural language…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Jecia Z. Y. Mao , Francis X. Creighton , Russell H. Taylor , Manish Sahu

We present SurgeonAssist-Net: a lightweight framework making action-and-workflow-driven virtual assistance, for a set of predefined surgical tasks, accessible to commercially available optical see-through head-mounted displays (OST-HMDs).…

Computer Vision and Pattern Recognition · Computer Science 2021-07-15 Mitchell Doughty , Karan Singh , Nilesh R. Ghugre

Real-time interactive video-chat portraits have been increasingly recognized as the future trend, particularly due to the remarkable progress made in text and voice chat technologies. However, existing methods primarily focus on real-time…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Jinwei Qi , Chaonan Ji , Sheng Xu , Peng Zhang , Bang Zhang , Liefeng Bo

Surgical videos captured from microscopic or endoscopic imaging devices are rich but complex sources of information, depicting different tools and anatomical structures utilized during an extended amount of time. Despite containing crucial…

Computer Vision and Pattern Recognition · Computer Science 2023-10-25 Felix Holm , Ghazal Ghazaei , Tobias Czempiel , Ege Özsoy , Stefan Saur , Nassir Navab

Modern surgeries are performed in complex and dynamic settings, including ever-changing interactions between medical staff, patients, and equipment. The holistic modeling of the operating room (OR) is, therefore, a challenging but essential…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Ege Özsoy , Tobias Czempiel , Felix Holm , Chantal Pellegrini , Nassir Navab

Automated surgical workflow analysis and understanding can assist surgeons to standardize procedures and enhance post-surgical assessment and indexing, as well as, interventional monitoring. Computer-assisted interventional (CAI) systems…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Odysseas Zisimopoulos , Evangello Flouty , Imanol Luengo , Petros Giataganas , Jean Nehme , Andre Chow , Danail Stoyanov

Minimally invasive surgery has dramatically improved patient operative outcomes, yet identifying safe operative zones remains challenging in critical phases, requiring surgeons to integrate visual cues, procedural phase, and anatomical…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Guanyi Qin , Xiaozhen Wang , Zhu Zhuo , Chang Han Low , Yuancan Xiao , Yibing Fu , Haofeng Liu , Kai Wang , Chunjiang Li , Yueming Jin

Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without relying on expert-labeled datasets. However, progress in…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Alejandra Perez , Chinedu Nwoye , Ramtin Raji Kermani , Omid Mohareri , Muhammad Abdullah Jamal

Multimodal Large Language Models (MLLMs) have achieved strong performance across many tasks, yet most systems remain limited to offline inference, requiring complete inputs before generating outputs. Recent streaming methods reduce latency…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Junyan Lin , Junlong Tong , Hao Wu , Jialiang Zhang , Jinming Liu , Xin Jin , Xiaoyu Shen

Multimodal Large Language Models have achieved significant success in offline video understanding, yet their application to streaming videos is severely limited by the linear explosion of visual tokens, which often leads to Out-of-Memory…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Chao Wang , Xudong Tan , Jianjian Cao , Kangcong Li , Tao Chen

Instructional video generation is an emerging task that aims to synthesize coherent demonstrations of procedural activities from textual descriptions. Such capability has broad implications for content creation, education, and human-AI…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Cheeun Hong , German Barquero , Fadime Sener , Markos Georgopoulos , Edgar Schönfeld , Stefan Popov , Yuming Du , Oscar Mañas , Albert Pumarola

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…

Reconstructing surgical scenes from monocular endoscopic video is critical for advancing robotic-assisted surgery. However, the application of state-of-the-art general-purpose reconstruction models is constrained by two key challenges: the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Kaiyuan Xu , Fangzhou Hong , Daniel Elson , Baoru Huang

Recent advancements in surgical computer vision applications have been driven by vision-only models, which do not explicitly integrate the rich semantics of language into their design. These methods rely on manually annotated surgical…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Kun Yuan , Vinkle Srivastav , Tong Yu , Joel L. Lavanchy , Jacques Marescaux , Pietro Mascagni , Nassir Navab , Nicolas Padoy

Accurate and robust tracking and reconstruction of the surgical scene is a critical enabling technology toward autonomous robotic surgery. Existing algorithms for 3D perception in surgery mainly rely on geometric information, while we…

Image and Video Processing · Electrical Eng. & Systems 2023-02-21 Shan Lin , Albert J. Miao , Jingpei Lu , Shunkai Yu , Zih-Yun Chiu , Florian Richter , Michael C. Yip

To develop intelligent speech assistants and integrate them seamlessly with intra-operative decision-support frameworks, accurate and efficient surgical phase recognition is a prerequisite. In this study, we propose a multimodal framework…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-24 Kubilay Can Demir , Belen Lojo Rodriguez , Tobias Weise , Andreas Maier , Seung Hee Yang

After decades of use in dictation and, more recently, ambient documentation, speech is emerging as a primary modality for interacting with technology and AI in healthcare. Yet medical speech recognition remains difficult: systems must…

Machine Learning · Computer Science 2026-05-22 Arne Nix , Robert James , Lasse Borgholt , Anna B. Ekner , Lana Krumm , Julius Severin , Dan Engel , Lars Maaløe , Jakob Havtorn