English
Related papers

Related papers: A generalizable foundation model for intraoperativ…

200 papers

Surgical scene perception via videos is critical for advancing robotic surgery, telesurgery, and AI-assisted surgery, particularly in ophthalmology. However, the scarcity of diverse and richly annotated video datasets has hindered the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Ming Hu , Peng Xia , Lin Wang , Siyuan Yan , Feilong Tang , Zhongxing Xu , Yimin Luo , Kaimin Song , Jurgen Leitner , Xuelian Cheng , Jun Cheng , Chi Liu , Kaijing Zhou , Zongyuan Ge

Surgical robotics is a rising field in medical technology and advanced robotics. Robot assisted surgery, or robotic surgery, allows surgeons to perform complicated surgical tasks with more precision, automation, and flexibility than is…

Robotics · Computer Science 2023-11-02 A. E. Abdelaal , S. N. Zaman , P. Y Chen , T. Suzuki , J. Ingleton

Image decomposition aims to analyze an image into elementary components, which is essential for numerous downstream tasks and also by nature provides certain interpretability to the analysis. Deep learning can be powerful for such tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Sihan Wang , Shangqi Gao , Fuping Wu , Xiahai Zhuang

Domain generalisation involves learning artificial intelligence (AI) models that can maintain high performance across diverse domains within a specific task. In video games, for instance, such AI models can supposedly learn to detect player…

Human-Computer Interaction · Computer Science 2024-09-23 Kosmas Pinitas , Konstantinos Makantasis , Georgios N. Yannakakis

Recording surgery in operating rooms is an essential task for education and evaluation of medical treatment. However, recording the desired targets, such as the surgery field, surgical tools, or doctor's hands, is difficult because the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Ryo Hachiuma , Tomohiro Shimizu , Hideo Saito , Hiroki Kajita , Yoshifumi Takatsume

With the advent of robot-assisted surgery, the role of data-driven approaches to integrate statistics and machine learning is growing rapidly with prominent interests in objective surgical skill assessment. However, most existing work…

Computer Vision and Pattern Recognition · Computer Science 2019-03-08 Ziheng Wang , Ann Majewicz Fey

Healthcare robotics requires robust multimodal perception and reasoning to ensure safety in dynamic clinical environments. Current Vision-Language Models (VLMs) demonstrate strong general-purpose capabilities but remain limited in temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Saurav Jha , Stefan K. Ehrlich

Automated surgical step recognition is an important task that can significantly improve patient safety and decision-making during surgeries. Existing state-of-the-art methods for surgical step recognition either rely on separate,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-22 Nisarg A. Shah , Shameema Sikder , S. Swaroop Vedula , Vishal M. Patel

We propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a single model. Most previous 3D segmentation approaches are…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Wei Xu , Chunsheng Shi , Sifan Tu , Xin Zhou , Dingkang Liang , Xiang Bai

A surgical world model capable of generating realistic surgical action videos with precise control over tool-tissue interactions can address fundamental challenges in surgical AI and simulation -- from data scarcity and rare event synthesis…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Sampath Rapuri , Lalithkumar Seenivasan , Dominik Schneider , Roger Soberanis-Mukul , Yufan He , Hao Ding , Jiru Xu , Chenhao Yu , Chenyan Jing , Pengfei Guo , Daguang Xu , Mathias Unberath

The success of large language models has inspired the computer vision community to explore image segmentation foundation model that is able to zero/few-shot generalize through prompt engineering. Segment-Anything(SAM), among others, is the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Haojie Zhang , Yongyi Su , Xun Xu , Kui Jia

Recent advances in large-scale pretraining have yielded visual foundation models with strong capabilities. Not only can recent models generalize to arbitrary images for their training task, their intermediate representations are useful for…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Mohamed El Banani , Amit Raj , Kevis-Kokitsi Maninis , Abhishek Kar , Yuanzhen Li , Michael Rubinstein , Deqing Sun , Leonidas Guibas , Justin Johnson , Varun Jampani

Open, or non-laparoscopic surgery, represents the vast majority of all operating room procedures, but few tools exist to objectively evaluate these techniques at scale. Current efforts involve human expert-based visual assessment. We…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Michael Zhang , Xiaotian Cheng , Daniel Copeland , Arjun Desai , Melody Y. Guan , Gabriel A. Brat , Serena Yeung

Segment anything model (SAM) has demonstrated excellent generalizability in common vision scenarios, yet falling short of the ability to understand specialized data. Recently, several methods have combined parameter-efficient techniques…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Yiran Song , Qianyu Zhou , Xuequan Lu , Zhiwen Shao , Lizhuang Ma

Data-driven approaches to assist operating room (OR) workflow analysis depend on large curated datasets that are time consuming and expensive to collect. On the other hand, we see a recent paradigm shift from supervised learning to…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Muhammad Abdullah Jamal , Omid Mohareri

Disentangled representation learning has been proposed as an approach to learning general representations even in the absence of, or with limited, supervision. A good general representation can be fine-tuned for new target tasks using…

Computer Vision and Pattern Recognition · Computer Science 2022-08-01 Xiao Liu , Pedro Sanchez , Spyridon Thermos , Alison Q. O'Neil , Sotirios A. Tsaftaris

Computer-assisted surgery research requires large, deeply annotated video datasets that capture clinical and technical variability. Existing cataract surgery resources lack the diversity and annotation depth required to train generalizable…

We present ENSAM (Equivariant, Normalized, Segment Anything Model), a lightweight and promptable model for universal 3D medical image segmentation. ENSAM combines a SegResNet-based encoder with a prompt encoder and mask decoder in a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Elias Stenhede , Agnar Martin Bjørnstad , Arian Ranjbar

Multi-modal medical imaging enables comprehensive diagnostics, yet current foundation models process 2D (e.g. X-ray) and 3D (e.g. CT) data with separate, dimensionality-specific architectures. We present MultiMedVision, a unified framework…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Frank Li , Bardia Khosravi , Mohammadreza Chavoshi , Young Seok Jeon , Theo Dapamede , Hari Trivedi , Janice Newsome , Judy Gichoya

Medical image segmentation remains challenging due to the high cost of pixel-level annotations for training. In the context of weak supervision, clinician gaze data captures regions of diagnostic interest; however, its sparsity limits its…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Jingkun Chen , Haoran Duan , Xiao Zhang , Boyan Gao , Vicente Grau , Jungong Han