English
Related papers

Related papers: A generalizable 3D framework and model for self-su…

200 papers

2D visual foundation models, such as DINOv3, a self-supervised model trained on large-scale natural images, have demonstrated strong zero-shot generalization, capturing both rich global context and fine-grained structural cues. However, an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Yik San Cheng , Runkai Zhao , Weidong Cai

Transfer learning from natural image to medical image has established as one of the most practical paradigms in deep learning for medical image analysis. However, to fit this paradigm, 3D imaging tasks in the most prominent imaging…

Image and Video Processing · Electrical Eng. & Systems 2019-08-20 Zongwei Zhou , Vatsal Sodha , Md Mahfuzur Rahman Siddiquee , Ruibin Feng , Nima Tajbakhsh , Michael B. Gotway , Jianming Liang

Medical instrument segmentation in 3D ultrasound is essential for image-guided intervention. However, to train a successful deep neural network for instrument segmentation, a large number of labeled images are required, which is expensive…

Computer Vision and Pattern Recognition · Computer Science 2021-08-03 Hongxu Yang , Caifeng Shan , R. Arthur Bouwman , Lukas R. C. Dekker , Alexander F. Kolen , Peter H. N. de With

Vision Transformers (ViTs) excel in 3D medical segmentation but require massive annotated datasets. While Self-Supervised Learning (SSL) mitigates this using unlabeled data, it still faces strict privacy and logistical barriers.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Jiaqi Tang , Mengyan Zheng , Shu Zhang , Fandong Zhang , Qingchao Chen

Large-scale volumetric medical images with annotation are rare, costly, and time prohibitive to acquire. Self-supervised learning (SSL) offers a promising pre-training and feature extraction solution for many downstream tasks, as it only…

Computer Vision and Pattern Recognition · Computer Science 2023-03-16 Ke Yu , Li Sun , Junxiang Chen , Max Reynolds , Tigmanshu Chaudhary , Kayhan Batmanghelich

Deep learning models in medical image analysis often struggle with generalizability across domains and demographic groups due to data heterogeneity and scarcity. Traditional augmentation improves robustness, but fails under substantial…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Sebastian Doerrich , Francesco Di Salvo , Jonas Alle , Christian Ledig

Self-supervised learning (SSL) has recently achieved promising performance for 3D medical image analysis tasks. Most current methods follow existing SSL paradigm originally designed for photographic or natural images, which cannot…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Yankai Jiang , Mingze Sun , Heng Guo , Xiaoyu Bai , Ke Yan , Le Lu , Minfeng Xu

In this work, we present EndoDINO, a foundation model for GI endoscopy tasks that achieves strong generalizability by pre-training on a well-curated image dataset sampled from the largest known GI endoscopy video dataset in the literature.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Patrick Dermyer , Angad Kalra , Matt Schwartz

Recent advancements in foundation models, typically trained with self-supervised learning on large-scale and diverse datasets, have shown great potential in medical image analysis. However, due to the significant spatial heterogeneity of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-25 Lingxiao Luo , Xuanzhong Chen , Bingda Tang , Xinsheng Chen , Rong Han , Chengpeng Hu , Yujiang Li , Ting Chen

During raw-data acquisition in CT imaging, diverse factors can degrade the collected sinograms, with undersampling and noise leading to severe artifacts and noise in reconstructed images and compromising diagnostic accuracy. Conventional…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Xingyu Ai , Shaoyu Wang , Zhiyuan Jia , Ao Xu , Hongming Shan , Jianhua Ma , Qiegen Liu

In the paper, we present an approach for learning a single model that universally segments 33 anatomical structures, including vertebrae, pelvic bones, and abdominal organs. Our model building has to address the following challenges.…

Image and Video Processing · Electrical Eng. & Systems 2022-03-07 Pengbo Liu , Yang Deng , Ce Wang , Yuan Hui , Qian Li , Jun Li , Shiwei Luo , Mengke Sun , Quan Quan , Shuxin Yang , You Hao , Honghu Xiao , Chunpeng Zhao , Xinbao Wu , S. Kevin Zhou

Volume-wise labeling in 3D medical images is a time-consuming task that requires expertise. As a result, there is growing interest in using semi-supervised learning (SSL) techniques to train models with limited labeled data. However, the…

Image and Video Processing · Electrical Eng. & Systems 2023-10-18 Haonan Wang , Xiaomeng Li

Self-supervised learning (SSL) models have recently demonstrated remarkable performance across various tasks, including image segmentation. This study delves into the emergent characteristics of the Self-Distillation with No Labels (DINO)…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Joseph A. Gallego-Mejia , Anna Jungbluth , Laura Martínez-Ferrer , Matt Allen , Francisco Dorr , Freddie Kalaitzis , Raúl Ramos-Pollán

Understanding model decisions is crucial in medical imaging, where interpretability directly impacts clinical trust and adoption. Vision Transformers (ViTs) have demonstrated state-of-the-art performance in diagnostic imaging; however,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Leili Barekatain , Ben Glocker

Self-supervised pretrain techniques have been widely used to improve the downstream tasks' performance. However, real-world magnetic resonance (MR) studies usually consist of different sets of contrasts due to different acquisition…

Image and Video Processing · Electrical Eng. & Systems 2025-06-17 Badhan Kumar Das , Ajay Singh , Gengyan Zhao , Han Liu , Thomas J. Re , Dorin Comaniciu , Eli Gibson , Andreas Maier

Vision foundation models (VFMs) are pre-trained on extensive image datasets to learn general representations for diverse types of data. These models can subsequently be fine-tuned for specific downstream tasks, significantly boosting…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Shansong Wang , Mojtaba Safari , Qiang Li , Chih-Wei Chang , Richard LJ Qiu , Justin Roper , David S. Yu , Xiaofeng Yang

Transformers have demonstrated remarkable performance in natural language processing and computer vision. However, existing vision Transformers struggle to learn from limited medical data and are unable to generalize on diverse medical…

Image and Video Processing · Electrical Eng. & Systems 2023-04-06 Yunhe Gao , Mu Zhou , Di Liu , Zhennan Yan , Shaoting Zhang , Dimitris N. Metaxas

3D image segmentation plays an important role in biomedical image analysis. Many 2D and 3D deep learning models have achieved state-of-the-art segmentation performance on 3D biomedical image datasets. Yet, 2D and 3D models have their own…

Computer Vision and Pattern Recognition · Computer Science 2018-12-11 Hao Zheng , Yizhe Zhang , Lin Yang , Peixian Liang , Zhuo Zhao , Chaoli Wang , Danny Z. Chen

This paper demonstrates that spatial information can be used to learn interpretable representations in medical images using Self-Supervised Learning (SSL). Our proposed method, ISImed, is based on the observation that medical images exhibit…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Nabil Jabareen , Dongsheng Yuan , Sören Lukassen

AI Foundation models are gaining traction in various applications, including medical fields like radiology. However, medical foundation models are often tested on limited tasks, leaving their generalisability and biases unexplored. We…