English
Related papers

Related papers: Whether and When does Endoscopy Domain Pretraining…

200 papers

Pretraining on large natural image classification datasets such as ImageNet has aided model development on data-scarce 2D medical tasks. 3D medical tasks often have much less data than 2D medical tasks, prompting practitioners to rely on…

Image and Video Processing · Electrical Eng. & Systems 2023-04-04 Alexander Ke , Shih-Cheng Huang , Chloe P O'Connell , Michal Klimont , Serena Yeung , Pranav Rajpurkar

Accurate 3D scene reconstruction is essential for numerous medical tasks. Given the challenges in obtaining ground truth data, there has been an increasing focus on self-supervised learning (SSL) for endoscopic depth estimation as a basis…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Beilei Cui , Long Bai , Mobarakol Islam , An Wang , Zhiqi Ma , Yiming Huang , Feng Li , Zhen Chen , Zhongliang Jiang , Nassir Navab , Hongliang Ren

Learning visual representations of medical images (e.g., X-rays) is core to medical image understanding but its progress has been held back by the scarcity of human annotations. Existing work commonly relies on fine-tuning weights…

Computer Vision and Pattern Recognition · Computer Science 2022-09-21 Yuhao Zhang , Hang Jiang , Yasuhide Miura , Christopher D. Manning , Curtis P. Langlotz

Endoscopic surgery is the gold standard for robotic-assisted minimally invasive surgery, offering significant advantages in early disease detection and precise interventions. However, the complexity of surgical scenes, characterized by high…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Guankun Wang , Rui Tang , Mengya Xu , Long Bai , Huxin Gao , Hongliang Ren

Unsupervised video-based surgical instrument segmentation has the potential to accelerate the adoption of robot-assisted procedures by reducing the reliance on manual annotations. However, the generally low quality of optical flow in…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Yang Liu , Peiran Wu , Jiayu Huo , Gongyu Zhang , Zhen Yuan , Christos Bergeles , Rachel Sparks , Prokar Dasgupta , Alejandro Granados , Sebastien Ourselin

Medical image segmentation is a vital healthcare endeavor requiring precise and efficient models for appropriate diagnosis and treatment. Vision transformer (ViT)-based segmentation models have shown great performance in accomplishing this…

Computer Vision and Pattern Recognition · Computer Science 2023-08-03 Numan Saeed , Muhammad Ridzuan , Roba Al Majzoub , Mohammad Yaqub

Image segmentation has been increasingly applied in medical settings as recent developments have skyrocketed the potential applications of deep learning. Urology, specifically, is one field of medicine that is primed for the adoption of a…

Image and Video Processing · Electrical Eng. & Systems 2022-05-02 Zachary A Stoebner , Daiwei Lu , Seok Hee Hong , Nicholas L Kavoussi , Ipek Oguz

The field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a surgery: including which surgical steps are performed; and…

The scarcity of annotated medical images is a major bottleneck in developing learning models for medical image analysis. Hence, recent studies have focused on pretrained models with fewer annotation requirements that can be fine-tuned for…

Image and Video Processing · Electrical Eng. & Systems 2024-10-02 Jonghun Kim , Mansu Kim , Hyunjin Park

Vision Transformer (ViT) has become one of the most popular neural architectures due to its great scalability, computational efficiency, and compelling performance in many vision tasks. However, ViT has shown inferior performance to…

Computer Vision and Pattern Recognition · Computer Science 2022-10-25 Junfei Xiao , Yutong Bai , Alan Yuille , Zongwei Zhou

Optical Coherence Tomography (OCT) provides high-resolution cross-sectional images useful for diagnosing various diseases, but their distinct characteristics from natural images raise questions about whether large-scale pre-training on…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Zihao Han , Philippe De Wilde

Despite their impressive performance in various surgical scene understanding tasks, deep learning-based methods are frequently hindered from deploying to real-world surgical applications for various causes. Particularly, data collection,…

Image and Video Processing · Electrical Eng. & Systems 2023-06-29 An Wang , Mobarakol Islam , Mengya Xu , Hongliang Ren

Automatic video activity recognition is crucial across numerous domains like surveillance, healthcare, and robotics. However, recognizing human activities from video data becomes challenging when training and test data stem from diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Partho Ghosh , Raisa Bentay Hossain , Mohammad Zunaed , Taufiq Hasan

Video capsule endoscopy has transformed gastrointestinal endoscopy (GIE) diagnostics by offering a non-invasive method for capturing detailed images of the gastrointestinal tract, enabling early disease detection. However, its potential is…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Marcel Roth , Micha V. Nowak , Adrian Krenzer , Frank Puppe

Self-supervised learning has emerged as a powerful paradigm for training deep neural networks, particularly in medical imaging where labeled data is scarce. While current approaches typically rely on synthetic augmentations of single…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Andre Dourson , Kylie Taylor , Xiaoli Qiao , Michael Fitzke

Masked Autoencoder (MAE) has recently been shown to be effective in pre-training Vision Transformers (ViT) for natural image analysis. By reconstructing full images from partially masked inputs, a ViT encoder aggregates contextual…

Image and Video Processing · Electrical Eng. & Systems 2023-04-24 Lei Zhou , Huidong Liu , Joseph Bae , Junjun He , Dimitris Samaras , Prateek Prasanna

Annotation of medical images has been a major bottleneck for the development of accurate and robust machine learning models. Annotation is costly and time-consuming and typically requires expert knowledge, especially in the medical domain.…

Computer Vision and Pattern Recognition · Computer Science 2020-01-08 Holger Roth , Ling Zhang , Dong Yang , Fausto Milletari , Ziyue Xu , Xiaosong Wang , Daguang Xu

Robotic surgery has been proven to offer clear advantages during surgical procedures, however, one of the major limitations is obtaining haptic feedback. Since it is often challenging to devise a hardware solution with accurate force…

Computer Vision and Pattern Recognition · Computer Science 2019-04-02 Cong Gao , Xingtong Liu , Michael Peven , Mathias Unberath , Austin Reiter

The ability to quickly annotate medical imaging data plays a critical role in training deep learning frameworks for segmentation. Doing so for image volumes or video sequences is even more pressing as annotating these is particularly…

Computer Vision and Pattern Recognition · Computer Science 2021-07-20 Laurent Lejeune , Raphael Sznitman

Medical image annotation is a major hurdle for developing precise and robust machine learning models. Annotation is expensive, time-consuming, and often requires expert knowledge, particularly in the medical field. Here, we suggest using…

Computer Vision and Pattern Recognition · Computer Science 2020-09-28 Holger R Roth , Dong Yang , Ziyue Xu , Xiaosong Wang , Daguang Xu