English
Related papers

Related papers: Large-scale Self-supervised Video Foundation Model…

200 papers

The difficulty of extracting deep features from EEG data and effectively integrating information from multiple views presents significant challenges for developing a generalizable pretraining framework for EEG representation learning.…

Machine Learning · Computer Science 2025-06-23 Puchun Liu , C. L. Philip Chen , Yubin He , Tong Zhang

Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but surgical benchmarks in particular are often missing from prominent medical benchmark suites. Since…

Estimating the remaining surgery duration (RSD) during surgical procedures can be useful for OR planning and anesthesia dose estimation. With the recent success of deep learning-based methods in computer vision, several neural network…

Computer Vision and Pattern Recognition · Computer Science 2020-02-27 Dominik Rivoir , Sebastian Bodenstedt , Felix von Bechtolsheim , Marius Distler , Jürgen Weitz , Stefanie Speidel

Recently, large-scale pre-trained language-image models like CLIP have shown extraordinary capabilities for understanding spatial contents, but naively transferring such models to video recognition still suffers from unsatisfactory temporal…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Zhiwu Qing , Shiwei Zhang , Ziyuan Huang , Yingya Zhang , Changxin Gao , Deli Zhao , Nong Sang

An accurate detection and tracking of devices such as guiding catheters in live X-ray image acquisitions is an essential prerequisite for endovascular cardiac interventions. This information is leveraged for procedural guidance, e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Saahil Islam , Venkatesh N. Murthy , Dominik Neumann , Badhan Kumar Das , Puneet Sharma , Andreas Maier , Dorin Comaniciu , Florin C. Ghesu

Conversation agents powered by large language models are revolutionizing the way we interact with visual data. Recently, large vision-language models (LVLMs) have been extensively studied for both images and videos. However, these studies…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Juseong Jin , Chang Wook Jeong

Medical visual question answering (VQA) bridges the gap between visual information and clinical decision-making, enabling doctors to extract understanding from clinical images and videos. In particular, surgical VQA can enhance the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Long Bai , Guankun Wang , Mobarakol Islam , Lalithkumar Seenivasan , An Wang , Hongliang Ren

Surgical video understanding is pivotal for enabling automated intraoperative decision-making, skill assessment, and postoperative quality improvement. However, progress in developing surgical video foundation models (FMs) remains hindered…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Jianhui Wei , Zikai Xiao , Danyu Sun , Luqi Gong , Zongxin Yang , Zuozhu Liu , Jian Wu

Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation emerging as an important contributing factor. In fact, existing validation…

Other Quantitative Biology · Quantitative Biology 2026-04-22 Annika Reinke , Ziying O. Li , Minu D. Tizabi , Pascaline André , Marcel Knopp , Mika M. Rother , Ines P. Machado , Maria S. Altieri , Deepak Alapatt , Sophia Bano , Sebastian Bodenstedt , Oliver Burgert , Elvis C. S. Chen , Justin W. Collins , Olivier Colliot , Evangelia Christodoulou , Tobias Czempiel , Adrito Das , Reuben Docea , Daniel Donoho , Qi Dou , Jennifer Eckhoff , Sandy Engelhardt , Gabor Fichtinger , Philipp Fuernstahl , Pablo García Kilroy , Stamatia Giannarou , Stephen Gilbert , Ines Gockel , Patrick Godau , Jan Gödeke , Teodor P. Grantcharov , Tamas Haidegger , Alexander Hann , Makoto Hashizume , Charles Heitz , Rebecca Hisey , Hanna Hoffmann , Arnaud Huaulmé , Paul F. Jäger , Pierre Jannin , Anthony Jarc , Rohit Jena , Yueming Jin , Leo Joskowicz , Luc Joyeux , Max Kirchner , Axel Krieger , Gernot Kronreif , Kyle Lam , Shlomi Laufer , Joël L. Lavanchy , Gyusung I. Lee , Robert Lim , Peng Liu , Hani J. Marcus , Pietro Mascagni , Ozanan R. Meireles , Beat P. Mueller , Lars Mündermann , Hirenkumar Nakawala , Nassir Navab , Abdourahmane Ndong , Juliane Neumann , Felix Nickel , Marco Nolden , Chinedu Nwoye , Namkee Oh , Nicolas Padoy , Thomas Pausch , Micha Pfeiffer , Tim Rädsch , Hongliang Ren , Nicola Rieke , Dominik Rivoir , Duygu Sarikaya , Samuel Schmidgall , Matthias Seibold , Silvia Seidlitz , Alexander Seitel , Lalith Sharan , Jeffrey H. Siewerdsen , Vinkle Srivastav , Raphael Sznitman , Russell Taylor , Thuy N. Tran , Matthias Unberath , Fons van der Sommen , Martin Wagner , Amine Yamlahi , Shaohua K. Zhou , Aneeq Zia , Amin Madani , Danail Stoyanov , Stefanie Speidel , Daniel A. Hashimoto , Fiona R. Kolbinger , Lena Maier-Hein

Online surgical phase recognition plays a significant role towards building contextual tools that could quantify performance and oversee the execution of surgical workflows. Current approaches are limited since they train spatial feature…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Yang Liu , Maxence Boels , Luis C. Garcia-Peraza-Herrera , Tom Vercauteren , Prokar Dasgupta , Alejandro Granados , Sebastien Ourselin

Medical vision-language pre-training methods mainly leverage the correspondence between paired medical images and radiological reports. Although multi-view spatial images and temporal sequences of image-report pairs are available in…

Artificial Intelligence · Computer Science 2024-05-31 Jinxia Yang , Bing Su , Wayne Xin Zhao , Ji-Rong Wen

Self-supervised learning has witnessed great progress in vision and NLP; recently, it also attracted much attention to various medical imaging modalities such as X-ray, CT, and MRI. Existing methods mostly focus on building new pretext…

Computer Vision and Pattern Recognition · Computer Science 2022-09-27 Xinpeng Ding , Ziwei Liu , Xiaomeng Li

Surgical phase recognition from video is a technology that automatically classifies the progress of a surgical procedure and has a wide range of potential applications, including real-time surgical support, optimization of medical…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Satoshi Kondo

Pathological image analysis is a crucial field in computer-aided diagnosis, where deep learning is widely applied. Transfer learning using pre-trained models initialized on natural images has effectively improved the downstream pathological…

Image and Video Processing · Electrical Eng. & Systems 2023-10-30 Nan Ying , Yanli Lei , Tianyi Zhang , Shangqing Lyu , Chunhui Li , Sicheng Chen , Zeyu Liu , Yu Zhao , Guanglei Zhang

Owing to recent advances in machine learning and the ability to harvest large amounts of data during robotic-assisted surgeries, surgical data science is ripe for foundational work. We present a large dataset of surgical videos and their…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Aneeq Zia , Max Berniker , Rogerio Nespolo , Xiaorui Zhang , Conor Perreault , Ziheng Wang , Benjamin Mueller , Ryan Schmidt , Kiran Bhattacharyya , Xi Liu , Anthony Jarc

Reconstructing surgical scenes from monocular endoscopic video is critical for advancing robotic-assisted surgery. However, the application of state-of-the-art general-purpose reconstruction models is constrained by two key challenges: the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Kaiyuan Xu , Fangzhou Hong , Daniel Elson , Baoru Huang

Deep-Learning-based video recognition has shown promising improvements along with the development of large-scale datasets and spatiotemporal network architectures. In image recognition, learning spatially invariant features is a key factor…

Computer Vision and Pattern Recognition · Computer Science 2020-08-14 Taeoh Kim , Hyeongmin Lee , MyeongAh Cho , Ho Seong Lee , Dong Heon Cho , Sangyoun Lee

Vision-Language Models (VLMs) have shown significant potential in surgical scene analysis, yet existing models are limited by frame-level datasets and lack high-quality video data with procedural surgical knowledge. To address these…

Other Quantitative Biology · Quantitative Biology 2026-01-21 Yaoqian Li , Xikai Yang , Dunyuan Xu , Yang Yu , Litao Zhao , Xiaowei Hu , Jinpeng Li , Pheng-Ann Heng

Five billion people in the world lack access to quality surgical care. Surgeon skill varies dramatically, and many surgical patients suffer complications and avoidable harm. Improving surgical training and feedback would help to reduce the…

Computer Vision and Pattern Recognition · Computer Science 2018-07-25 Amy Jin , Serena Yeung , Jeffrey Jopling , Jonathan Krause , Dan Azagury , Arnold Milstein , Li Fei-Fei

Accurate surgical phase recognition is essential for analyzing procedural workflows, supporting intraoperative decision-making, and enabling data-driven improvements in surgical education and performance evaluation. In this work, we present…