中文
相关论文

相关论文: A benchmark for video-based laparoscopic skill ana…

200 篇论文

Vision-Language Models (VLMs) have shown significant potential in surgical scene analysis, yet existing models are limited by frame-level datasets and lack high-quality video data with procedural surgical knowledge. To address these…

其他定量生物学 · 定量生物学 2026-01-21 Yaoqian Li , Xikai Yang , Dunyuan Xu , Yang Yu , Litao Zhao , Xiaowei Hu , Jinpeng Li , Pheng-Ann Heng

Automated assessment of surgical skills using artificial intelligence (AI) provides trainees with instantaneous feedback. After bimanual tool motions are captured, derived kinematic metrics are reliable predictors of performance in…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Shekhar Madhav Khairnar , Huu Phong Nguyen , Alexis Desir , Carla Holcomb , Daniel J. Scott , Ganesh Sankaranarayanan

Skill assessment in procedural videos is crucial for the objective evaluation of human performance in settings such as manufacturing and procedural daily tasks. Current research on skill assessment has predominantly focused on sports and…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Michele Mazzamuto , Daniele Di Mauro , Gianpiero Francesca , Giovanni Maria Farinella , Antonino Furnari

Autonomous laparoscopic camera control must maintain a stable and safe surgical view under rapid tool-tissue interactions while remaining interpretable to surgeons. We present a strategy-grounded framework that couples high-level…

机器人学 · 计算机科学 2026-02-25 Keyu Zhou , Peisen Xu , Yahao Wu , Jiming Chen , Gaofeng Li , Shunlei Li

We present a method for assessing skill from video, applicable to a variety of tasks, ranging from surgery to drawing and rolling pizza dough. We formulate the problem as pairwise (who's better?) and overall (who's best?) ranking of video…

计算机视觉与模式识别 · 计算机科学 2018-03-30 Hazel Doughty , Dima Damen , Walterio Mayol-Cuevas

Contemporary artificial neural networks (ANN) are trained end-to-end, jointly learning both features and classifiers for the task of interest. Though enormously effective, this paradigm imposes significant costs in assembling annotated…

Surgical tool segmentation in endoscopic videos is an important component of computer assisted interventions systems. Recent success of image-based solutions using fully-supervised deep learning approaches can be attributed to the…

计算机视觉与模式识别 · 计算机科学 2020-07-23 Manish Sahu , Ronja Strömsdörfer , Anirban Mukhopadhyay , Stefan Zachow

Endoscopic depth estimation is a critical technology for improving the safety and precision of minimally invasive surgery. It has attracted considerable attention from researchers in medical imaging, computer vision, and robotics. Over the…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Ke Niu , Zeyun Liu , Xue Feng , Heng Li , Qika Lin , Kaize Shi

In the field of robotic manipulation, deep imitation learning is recognized as a promising approach for acquiring manipulation skills. Additionally, learning from diverse robot datasets is considered a viable method to achieve versatility…

机器人学 · 计算机科学 2024-03-20 Heecheol Kim , Yoshiyuki Ohmura , Yasuo Kuniyoshi

Video is an essential imaging modality for diagnostics, e.g. in ultrasound imaging, for endoscopy, or movement assessment. However, video hasn't received a lot of attention in the medical image analysis community. In the clinical practice,…

计算机视觉与模式识别 · 计算机科学 2020-05-20 Tianrui Liu , Qingjie Meng , Athanasios Vlontzos , Jeremy Tan , Daniel Rueckert , Bernhard Kainz

Over the past few years, surgical data science has attracted substantial interest from the machine learning (ML) community. Various studies have demonstrated the efficacy of emerging ML techniques in analysing surgical data, particularly…

图像与视频处理 · 电气工程与系统科学 2023-07-06 Adnan Qayyum , Hassan Ali , Massimo Caputo , Hunaid Vohra , Taofeek Akinosho , Sofiat Abioye , Ilhem Berrou , Paweł Capik , Junaid Qadir , Muhammad Bilal

Markerless motion capture has become an active field of research in computer vision in recent years. Its extensive applications are known in a great variety of fields, including computer animation, human motion analysis, biomedical…

计算机视觉与模式识别 · 计算机科学 2022-01-10 Doan Duy Vo , Russell Butler

Multimodal Large Language Models (MLLMs) have shown significant potential in surgical video understanding. With improved zero-shot performance and more effective human-machine interaction, they provide a strong foundation for advancing…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Ziyang Song , Zelin Zang , Xiaofan Ye , Boqiang Xu , Long Bai , Jinlin Wu , Hongliang Ren , Hongbin Liu , Jiebo Luo , Zhen Lei

Consensus amongst researchers and industry points to a lack of large, representative annotated datasets as the biggest obstacle to progress in the field of surgical data science. Advances in Self-Supervised Learning (SSL) represent a…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Deepak Alapatt , Aditya Murali , Vinkle Srivastav , Pietro Mascagni , AI4SafeChole Consortium , Nicolas Padoy

Self-supervised learning has emerged as a powerful paradigm for label-free model pretraining, particularly in the video domain, where manual annotation is costly and time-intensive. However, existing self-supervised approaches employ…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Akash Kumar , Ashlesha Kumar , Vibhav Vineet , Yogesh S Rawat

In this work, we take aim towards increasing the effectiveness of surgical assistant robots. We intended to make assistant robots safer by making them aware about the actions of surgeon, so it can take appropriate assisting actions. In…

Automated surgical workflow analysis and understanding can assist surgeons to standardize procedures and enhance post-surgical assessment and indexing, as well as, interventional monitoring. Computer-assisted interventional (CAI) systems…

计算机视觉与模式识别 · 计算机科学 2018-07-30 Odysseas Zisimopoulos , Evangello Flouty , Imanol Luengo , Petros Giataganas , Jean Nehme , Andre Chow , Danail Stoyanov

Automatically detecting vital signs in videos, such as the estimation of heart and respiration rates, is a challenging research problem in computer vision with important applications in the medical field. One of the key difficulties in…

计算机视觉与模式识别 · 计算机科学 2020-04-27 Florin Condrea , Victor-Andrei Ivan , Marius Leordeanu

Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, existing methods are forced to choose between small,…

机器人学 · 计算机科学 2026-02-26 Hao Luo , Ye Wang , Wanpeng Zhang , Haoqi Yuan , Yicheng Feng , Haiweng Xu , Sipeng Zheng , Zongqing Lu

While there has been remarkable progress in the performance of visual recognition algorithms, the state-of-the-art models tend to be exceptionally data-hungry. Large labeled training datasets, expensive and tedious to produce, are required…

计算机视觉与模式识别 · 计算机科学 2016-06-07 Fisher Yu , Ari Seff , Yinda Zhang , Shuran Song , Thomas Funkhouser , Jianxiong Xiao