English
Related papers

Related papers: Pre-training on High Definition X-ray Images: An E…

200 papers

In the advent of a digital health revolution, vast amounts of clinical data are being generated, stored and processed on a daily basis. This has made the storage and retrieval of large volumes of health-care data, especially,…

Computer Vision and Pattern Recognition · Computer Science 2019-05-10 Asif Shahriyar Sushmit , Shakib Uz Zaman , Ahmed Imtiaz Humayun , Taufiq Hasan , Mohammed Imamul Hassan Bhuiyan

Foundation models or pre-trained models have substantially improved the performance of various language, vision, and vision-language understanding tasks. However, existing foundation models can only perform the best in one type of tasks,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Xinsong Zhang , Yan Zeng , Jipeng Zhang , Hang Li

Existing Masked Image Modeling methods apply fixed mask patterns to guide the self-supervised training. As those mask patterns resort to different criteria to depict image contents, sticking to a fixed pattern leads to a limited vision cues…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Zhanzhou Feng , Shiliang Zhang

Existing deep learning methods for remote sensing image fusion often suffer from poor generalization when applied to unseen datasets due to the limited availability of real training data and the domain gap between different satellite…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Yongchuan Cui , Peng Liu , Yi Zeng

Convolutional Neural Networks have reached extremely high performances on the Face Recognition task. Largely used datasets, such as VGGFace2, focus on gender, pose and age variations trying to balance them to achieve better results.…

Computer Vision and Pattern Recognition · Computer Science 2020-11-23 Fabio Valerio Massoli , Giuseppe Amato , Fabrizio Falchi

Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlooking the complex 3D geometry inherent to surgical…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 John J. Han , Adam Schmidt , Muhammad Abdullah Jamal , Chinedu Nwoye , Anita Rau , Jie Ying Wu , Omid Mohareri

Self-supervised learning has emerged as a powerful approach for leveraging large-scale unlabeled data to improve model performance in various domains. In this paper, we explore masked self-supervised pre-training for text recognition…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Martin Kišš , Michal Hradiš

The diagnosis and treatment of chest diseases play a crucial role in maintaining human health. X-ray examination has become the most common clinical examination means due to its efficiency and cost-effectiveness. Artificial intelligence…

Computer Vision and Pattern Recognition · Computer Science 2024-05-09 Jingfeng Yao , Xinggang Wang , Yuehao Song , Huangxuan Zhao , Jun Ma , Yajie Chen , Wenyu Liu , Bo Wang

Foundation models have transformed vision and language by learning general-purpose representations from large-scale unlabeled data, yet 3D medical imaging lacks analogous approaches. Existing self-supervised methods rely on low-level…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yunhe Gao , Yabin Zhang , Chong Wang , Jiaming Liu , Maya Varma , Jean-Benoit Delbrouck , Akshay Chaudhari , Curtis Langlotz

Pre-training has marked numerous state of the arts in high-level computer vision, while few attempts have ever been made to investigate how pre-training acts in image processing systems. In this paper, we tailor transformer-based…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Wenbo Li , Xin Lu , Shengju Qian , Jiangbo Lu , Xiangyu Zhang , Jiaya Jia

Obtaining large pre-trained models that can be fine-tuned to new tasks with limited annotated samples has remained an open challenge for medical imaging data. While pre-trained deep networks on ImageNet and vision-language foundation models…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Duy M. H. Nguyen , Hoang Nguyen , Nghiem T. Diep , Tan N. Pham , Tri Cao , Binh T. Nguyen , Paul Swoboda , Nhat Ho , Shadi Albarqouni , Pengtao Xie , Daniel Sonntag , Mathias Niepert

Strong gravitational lensing can reveal the influence of dark-matter substructure in galaxies, but analyzing these effects from noisy, low-resolution images poses a significant challenge. In this work, we propose a masked autoencoder (MAE)…

Harnessing the power of pre-training on large-scale datasets like ImageNet forms a fundamental building block for the progress of representation learning-driven solutions in computer vision. Medical images are inherently different from…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Jeya Maria Jose Valanarasu , Yucheng Tang , Dong Yang , Ziyue Xu , Can Zhao , Wenqi Li , Vishal M. Patel , Bennett Landman , Daguang Xu , Yufan He , Vishwesh Nath

Deep learning methods for chest X-ray interpretation typically rely on pretrained models developed for ImageNet. This paradigm assumes that better ImageNet architectures perform better on chest X-ray tasks and that ImageNet-pretrained…

Computer Vision and Pattern Recognition · Computer Science 2021-02-23 Alexander Ke , William Ellsworth , Oishi Banerjee , Andrew Y. Ng , Pranav Rajpurkar

Self-supervised foundation models have shown great potential in computer vision thanks to the pre-training paradigm of masked autoencoding. Scale is a primary factor influencing the performance of these foundation models. However, these…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Zhiyu Zhao , Bingkun Huang , Sen Xing , Gangshan Wu , Yu Qiao , Limin Wang

Advancements in diffusion-based foundation models have improved text-to-image generation, yet most efforts have been limited to low-resolution settings. As high-resolution image synthesis becomes increasingly essential for various…

Image and Video Processing · Electrical Eng. & Systems 2025-08-22 Zahra TehraniNasab , Amar Kumar , Tal Arbel

The complex structure of the heart leads to significant challenges in echocardiography, especially in acquisition cardiac ultrasound images. Successful echocardiography requires a thorough understanding of the structures on the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Haojun Jiang , Meng Li , Zhenguo Sun , Ning Jia , Yu Sun , Shaqi Luo , Shiji Song , Gao Huang

In this work, we explore self-supervised visual pre-training on images from diverse, in-the-wild videos for real-world robotic tasks. Like prior work, our visual representations are pre-trained via a masked autoencoder (MAE), frozen, and…

Robotics · Computer Science 2022-10-07 Ilija Radosavovic , Tete Xiao , Stephen James , Pieter Abbeel , Jitendra Malik , Trevor Darrell

Contrastive language-image pretraining has shown great success in learning visual-textual joint representation from web-scale data, demonstrating remarkable "zero-shot" generalization ability for various image tasks. However, how to…

Computer Vision and Pattern Recognition · Computer Science 2022-08-05 Bolin Ni , Houwen Peng , Minghao Chen , Songyang Zhang , Gaofeng Meng , Jianlong Fu , Shiming Xiang , Haibin Ling

Foundation models leverage large-scale pretraining to capture extensive knowledge, demonstrating generalization in a wide range of language tasks. By comparison, vision foundation models (VFMs) often exhibit uneven improvements across…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Shiqi Huang , Yipei Wang , Natasha Thorley , Alexander Ng , Shaheer Saeed , Mark Emberton , Shonit Punwani , Veeru Kasivisvanathan , Dean Barratt , Daniel Alexander , Yipeng Hu