English
Related papers

Related papers: MOSaiC: a Web-based Platform for Collaborative Med…

200 papers

In this work, we take aim towards increasing the effectiveness of surgical assistant robots. We intended to make assistant robots safer by making them aware about the actions of surgeon, so it can take appropriate assisting actions. In…

Enabled by large annotated datasets, tracking and segmentation of objects in videos has made remarkable progress in recent years. Despite these advancements, algorithms still struggle under degraded conditions and during fast movements.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Friedhelm Hamann , Hanxiong Li , Paul Mieske , Lars Lewejohann , Guillermo Gallego

To meet the growing demand for systematic surgical training, wet-lab environments have become indispensable platforms for hands-on practice in ophthalmology. Yet, traditional wet-lab training depends heavily on manual performance…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Negin Ghamsarian , Raphael Sznitman , Klaus Schoeffmann , Jens Kowal

Localisation of surgical tools constitutes a foundational building block for computer-assisted interventional technologies. Works in this field typically focus on training deep learning models to perform segmentation tasks. Performance of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Zhe Han , Charlie Budd , Gongyu Zhang , Huanyu Tian , Christos Bergeles , Tom Vercauteren

Machine learning is transforming the video editing industry. Recent advances in computer vision have leveled-up video editing tasks such as intelligent reframing, rotoscoping, color grading, or applying digital makeups. However, most of the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Dawit Mureja Argaw , Fabian Caba Heilbron , Joon-Young Lee , Markus Woodson , In So Kweon

We present 3MASSIV, a multilingual, multimodal and multi-aspect, expertly-annotated dataset of diverse short videos extracted from short-video social media platform - Moj. 3MASSIV comprises of 50k short videos (20 seconds average duration)…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Vikram Gupta , Trisha Mittal , Puneet Mathur , Vaibhav Mishra , Mayank Maheshwari , Aniket Bera , Debdoot Mukherjee , Dinesh Manocha

The surgical operating room (OR) presents many opportunities for automation and optimization. Videos from various sources in the OR are becoming increasingly available. The medical community seeks to leverage this wealth of data to develop…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Lennart Bastian , Tobias Czempiel , Christian Heiliger , Konrad Karcz , Ulrich Eck , Benjamin Busam , Nassir Navab

Audio-visual learning seeks to enhance the computer's multi-modal perception leveraging the correlation between the auditory and visual modalities. Despite their many useful downstream tasks, such as video retrieval, AR/VR, and…

Human-Computer Interaction · Computer Science 2023-07-31 Zheng Zhang , Zheng Ning , Chenliang Xu , Yapeng Tian , Toby Jia-Jun Li

We strive for spatio-temporal localization of actions in videos. The state-of-the-art relies on action proposals at test time and selects the best one with a classifier trained on carefully annotated box annotations. Annotating action boxes…

Computer Vision and Pattern Recognition · Computer Science 2017-12-14 Pascal Mettes , Jan C. van Gemert , Cees G. M. Snoek

Video object segmentation is an emerging technology that is well-suited for real-time surgical video segmentation, offering valuable clinical assistance in the operating room by ensuring consistent frame tracking. However, its adoption is…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Seyed Amir Mousavi , Utku Ozbulak , Francesca Tozzi , Nikdokht Rashidian , Wouter Willaert , Joris Vankerschaver , Wesley De Neve

We introduce a dataset for evidence/rationale extraction on an extreme multi-label classification task over long medical documents. One such task is Computer-Assisted Coding (CAC) which has improved significantly in recent years, thanks to…

Computation and Language · Computer Science 2023-07-11 Hua Cheng , Rana Jafari , April Russell , Russell Klopfer , Edmond Lu , Benjamin Striner , Matthew R. Gormley

This chapter describes how astronomical imaging survey data have become a vital part of modern astronomy, how these data are archived and then served to the astronomical community through on-line data access portals. The Virtual…

Instrumentation and Methods for Astrophysics · Physics 2015-03-17 Daniel S. Katz , G. Bruce Berriman , Robert G. Mann

Digital pathology plays a crucial role in the development of artificial intelligence in the medical field. The digital pathology platform can make the pathological resources digital and networked, and realize the permanent storage of visual…

Human-Computer Interaction · Computer Science 2021-11-11 Jialun Wu , Anyu Mao , Xinrui Bao , Haichuan Zhang , Zeyu Gao , Chunbao Wang , Tieliang Gong , Chen Li

Video mosaicking requires the registration of overlapping frames located at distant timepoints in the sequence to ensure global consistency of the reconstructed scene. However, fully automated registration of such long-range pairs is (i)…

Computer Vision and Pattern Recognition · Computer Science 2021-01-01 Loic Peter , Marcel Tella-Amo , Dzhoshkun Ismail Shakir , Jan Deprest , Sebastien Ourselin , Juan Eugenio Iglesias , Tom Vercauteren

Users often take notes for instructional videos to access key knowledge later without revisiting long videos. Automated note generation tools enable users to obtain informative notes efficiently. However, notes generated by existing…

Human-Computer Interaction · Computer Science 2025-08-21 Running Zhao , Zhihan Jiang , Xinchen Zhang , Chirui Chang , Handi Chen , Weipeng Deng , Luyao Jin , Xiaojuan Qi , Xun Qian , Edith C. H. Ngai

Active learning (AL) can reduce annotation costs in surgical video analysis while maintaining model performance. However, traditional AL methods, developed for images or short video clips, are suboptimal for surgical step recognition due to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Nisarg A. Shah , Bardia Safaei , Shameema Sikder , S. Swaroop Vedula , Vishal M. Patel

Purpose: In medical research, deep learning models rely on high-quality annotated data, a process often laborious and timeconsuming. This is particularly true for detection tasks where bounding box annotations are required. The need to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Meyer Adrien , Mazellier Jean-Paul , Jeremy Dana , Nicolas Padoy

We present CataractSAM-2, a domain-adapted extension of Meta's Segment Anything Model 2, designed for real-time semantic segmentation of cataract ophthalmic surgery videos with high accuracy. Positioned at the intersection of computer…

Few-shot video object segmentation aims to reduce annotation costs; however, existing methods still require abundant dense frame annotations for training, which are scarce in the medical domain. We investigate an extremely low-data regime…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Zixuan Zheng , Yilei Shi , Chunlei Li , Jingliang Hu , Xiao Xiang Zhu , Lichao Mou