English
Related papers

Related papers: Large-scale Self-supervised Video Foundation Model…

200 papers

Self-supervised tasks have been utilized to build useful representations that can be used in downstream tasks when the annotation is unavailable. In this paper, we introduce a self-supervised video representation learning method based on…

Computer Vision and Pattern Recognition · Computer Science 2021-02-23 Duc Quang Vu , Ngan T. H. Le , Jia-Ching Wang

Recent self-supervised advances in medical computer vision exploit global and local anatomical self-similarity for pretraining prior to downstream tasks such as segmentation. However, current methods assume i.i.d. image acquisition, which…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Mengwei Ren , Neel Dey , Martin A. Styner , Kelly Botteron , Guido Gerig

Automated, clinician-grade assessment reports for surgical procedures could reduce documentation burden and provide objective feedback, yet remain challenging due to the difficulty of aligning dense spatio-temporal video representations…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Kedi Sun , Chaohui Dang , Yue Feng , James Glasbey , Theodoros N. Arvanitis , Le Zhang

Cochlear Implant (CI) procedures involve performing an invasive mastoidectomy to insert an electrode array into the cochlea. In this paper, we introduce a novel pipeline that is capable of generating synthetic multi-view videos from a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Yike Zhang , Jack Noble

Minimally invasive image-guided surgery heavily relies on vision. Deep learning models for surgical video analysis could therefore support visual tasks such as assessing the critical view of safety (CVS) in laparoscopic cholecystectomy…

Image and Video Processing · Electrical Eng. & Systems 2021-09-21 Pietro Mascagni , Deepak Alapatt , Alain Garcia , Nariaki Okamoto , Armine Vardazaryan , Guido Costamagna , Bernard Dallemagne , Nicolas Padoy

Learning meaningful visual representations in an embedding space can facilitate generalization in downstream tasks such as action segmentation and imitation. In this paper, we learn a motion-centric representation of surgical video…

Robotics · Computer Science 2020-06-02 Ajay Kumar Tanwani , Pierre Sermanet , Andy Yan , Raghav Anand , Mariano Phielipp , Ken Goldberg

Video Large Language Models (Video-LLMs) require continual learning to adapt to non-stationary real-world data. However, existing benchmarks fall short of evaluating modern foundation models: many still rely on models without large-scale…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Haiyang Guo , Yichen Shi , Fei Zhu , Wenzhuo Liu , Hongbo Zhao , Fanhu Zeng , Shijie Ma , Da-Han Wang , Xu-Yao Zhang

Surgical phase recognition is a critical component for context-aware decision support in intelligent operating rooms, yet training robust models is hindered by limited annotated clinical videos and large domain gaps between synthetic and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Yuxin He , An Li , Cheng Xue

Robotic assisted (RA) surgery promises to transform surgical intervention. Intuitive Surgical is committed to fostering these changes and the machine learning models and algorithms that will enable them. With these goals in mind we have…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Aneeq Zia , Max Berniker , Rogerio Garcia Nespolo , Xiaorui Zhang , Conor Perreault , Kiran Bhattacharyya , Xi Liu , Ziheng Wang , Satoshi Kondo , Satoshi Kasai , Kousuke Hirasawa , Bo Liu , David Austin , Yiheng Wang , Michal Futrega , Jean-Francois Puget , Zhenqiang Li , Yoichi Sato , Ryo Fujii , Ryo Hachiuma , Mana Masuda , Hideo Saito , An Wang , Mengya Xu , Mobarakol Islam , Long Bai , Winnie Pang , Hongliang Ren , Chinedu Nwoye , Luca Sestini , Nicolas Padoy , Maximilian Nielsen , Samuel Schüttler , Thilo Sentker , Hümeyra Husseini , Ivo Baltruschat , Rüdiger Schmitz , René Werner , Aleksandr Matsun , Mugariya Farooq , Numan Saaed , Jose Renato Restom Viera , Mohammad Yaqub , Neil Getty , Fangfang Xia , Zixuan Zhao , Xiaotian Duan , Xing Yao , Ange Lou , Hao Yang , Jintong Han , Jack Noble , Jie Ying Wu , Tamer Abdulbaki Alshirbaji , Nour Aldeen Jalal , Herag Arabian , Ning Ding , Knut Moeller , Weiliang Chen , Quan He , Muhammad Bilal , Taofeek Akinosho , Adnan Qayyum , Massimo Caputo , Hunaid Vohra , Michael Loizou , Anuoluwapo Ajayi , Ilhem Berrou , Faatihah Niyi-Odumosu , Charlie Budd , Oluwatosin Alabi , Tom Vercauteren , Ruoxi Zhao , Ayberk Acar , John Han , Jumanh Atoum , Yinhong Qin , Surong Hua , Lu Ping , Wenming Wu , Rongfeng Wei , Jinlin Wu , You Pang , Zhen Chen , Tim Jaspers , Amine Yamlahi , Piotr Kalinowski , Dominik Michael , Tim Rädsch , Marco Hübner , Danail Stoyanov , Stefanie Speidel , Lena Maier-Hein , Jie Tian , Ruxin Zhang , Khang Hoang Nguyen , Anh Quoc Nguyen , Tam Minh Nguyen , Khoi Dinh Tran , Minh Nguyen Dang Nhat , Trinh Thi Doan Pham , Linh Van Nguyen , Chunyang Jiang , Dewei Yang , Haitao Li , Yannick Prudent , Thibaut Boissin , Mahmood Alam , Shazad Ashraf , Andrew D. Beggs , Lukman Akanbi , Manuel D. Delgado , Narain Gupta , Amir M. Hajiyavand , Iqbal Qasim , Hafiz A. Alaka , Junaid Qadir , Shu Yang , Yihui Wang , Hao Chen , Shin Paul , Yosuke Yamagishi , Zhang Dong , Hongyun Li , Hongyu Gu , Xiaoliu Ding , Xiaoyao Liu , Xingyu Zhao , Mariana Ribeiro , Tiago Jesus , André Ferreira , Guilherme Barbosa , João Carvalho , Leonardo Barroso , Nuno Gomes , Rafael Peixoto , Rodrigo Ralha , Victor Alves , Stephanie , Nattapat Ittikosil , Achita Chitrapan , Quan Huu Cap , Jiayuan Huang , Shreyas C Dhake , Sergi Kavtaradze , Mobarak I Hoque , Ka Young Kim , Su Yong Yun , Young Tae Kim , Hyeon Bae Kim , Seong Tae Kim , Zuxing Deng , Ling Li , Jieyu Zheng , Xiaojian Li , Anthony Jarc

Recently, spatiotemporal graphs have emerged as a concise and elegant manner of representing video clips in an object-centric fashion, and have shown to be useful for downstream tasks such as action recognition. In this work, we investigate…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Aditya Murali , Deepak Alapatt , Pietro Mascagni , Armine Vardazaryan , Alain Garcia , Nariaki Okamoto , Didier Mutter , Nicolas Padoy

Natural language could play an important role in developing generalist surgical models by providing a broad source of supervision from raw texts. This flexible form of supervision can enable the model's transferability across datasets and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Kun Yuan , Vinkle Srivastav , Nassir Navab , Nicolas Padoy

Real-time algorithms for automatically recognizing surgical phases are needed to develop systems that can provide assistance to surgeons, enable better management of operating room (OR) resources and consequently improve safety within the…

Computer Vision and Pattern Recognition · Computer Science 2018-05-23 Gaurav Yengera , Didier Mutter , Jacques Marescaux , Nicolas Padoy

Purpose: The objective of this investigation is to provide a comprehensive analysis of state-of-the-art methods for video-based assessment of surgical skill in the operating room. Methods: Using a data set of 99 videos of capsulorhexis, a…

Computer Vision and Pattern Recognition · Computer Science 2022-05-16 Sanchit Hira , Digvijay Singh , Tae Soo Kim , Shobhit Gupta , Gregory Hager , Shameema Sikder , S. Swaroop Vedula

Data-driven approaches to assist operating room (OR) workflow analysis depend on large curated datasets that are time consuming and expensive to collect. On the other hand, we see a recent paradigm shift from supervised learning to…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Muhammad Abdullah Jamal , Omid Mohareri

Artificial intelligence is set to be deployed in operating rooms to improve surgical care. This early-stage clinical evaluation shows the feasibility of concurrently attaining real-time, high-quality predictions from several deep neural…

Image and Video Processing · Electrical Eng. & Systems 2022-12-14 Pietro Mascagni , Deepak Alapatt , Alfonso Lapergola , Armine Vardazaryan , Jean-Paul Mazellier , Bernard Dallemagne , Didier Mutter , Nicolas Padoy

Over the past one hundred years, the classic teaching methodology of "see one, do one, teach one" has governed the surgical education systems worldwide. With the advent of Operation Room 2.0, recording video, kinematic and many other types…

Computer Vision and Pattern Recognition · Computer Science 2019-07-23 Hassan Ismail Fawaz , Germain Forestier , Jonathan Weber , François Petitjean , Lhassane Idoumghar , Pierre-Alain Muller

Unsupervised multi-object scene decomposition is a fast-emerging problem in representation learning. Despite significant progress in static scenes, such models are unable to leverage important dynamic cues present in video. We propose a…

Computer Vision and Pattern Recognition · Computer Science 2020-06-29 Polina Zablotskaia , Edoardo A. Dominici , Leonid Sigal , Andreas M. Lehrmann

Vision Transformer models have shown impressive effectiveness in the surgical video understanding tasks through long-range dependency modeling. However, current methods suffer from prohibitive computational costs due to processing massive…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Xixi Jiang , Chen Yang , Dong Zhang , Pingcheng Dong , Xin Yang , Kwang-Ting Cheng

Laparoscopic surgery is a complex surgical technique that requires extensive training. Recent advances in deep learning have shown promise in supporting this training by enabling automatic video-based assessment of surgical skills. However,…

Tissue manipulation is a frequently used fundamental subtask of any surgical procedures, and in some cases it may require the involvement of a surgeon's assistant. The complex dynamics of soft tissue as an unstructured environment is one of…

‹ Prev 1 3 4 5 6 7 10 Next ›