中文
相关论文

相关论文: InfMAE: A Foundation Model in the Infrared Modalit…

200 篇论文

In this work, we survey recent studies on masked image modeling (MIM), an approach that emerged as a powerful self-supervised learning technique in computer vision. The MIM task involves masking some information, e.g. pixels, patches, or…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Vlad Hondru , Florinel Alin Croitoru , Shervin Minaee , Radu Tudor Ionescu , Nicu Sebe

Artificial Intelligence (AI) technologies have profoundly transformed the field of remote sensing, revolutionizing data collection, processing, and analysis. Traditionally reliant on manual interpretation and task-specific models, remote…

计算机视觉与模式识别 · 计算机科学 2025-02-13 Siqi Lu , Junlin Guo , James R Zimmer-Dauphinee , Jordan M Nieusma , Xiao Wang , Parker VanValkenburgh , Steven A Wernke , Yuankai Huo

Clinical deployment of automated brain MRI analysis faces a fundamental challenge: clinical data is heterogeneous and noisy, and high-quality labels are prohibitively costly to obtain. Self-supervised learning (SSL) can address this by…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Asbjørn Munk , Stefano Cerri , Vardan Nersesjan , Christian Hedeager Krag , Jakob Ambsdorf , Pablo Rocamora García , Julia Machnio , Peirong Liu , Suhyun Ahn , Nasrin Akbari , Yasmina Al Khalil , Kimberly Amador , Sina Amirrajab , Tal Arbel , Meritxell Bach Cuadra , Ujjwal Baid , Bhakti Baheti , Jaume Banus , Kamil Barbierik , Christoph Brune , Yansong Bu , Baptiste Callard , Yuhan Chen , Cornelius Crijnen , Corentin Dancette , Peter Drotar , Prasad Dutande , Nils D. Forkert , Saurabh Garg , Jakub Gazda , Matej Gazda , Benoît Gérin , Partha Ghosh , Weikang Gong , Pedro M. Gordaliza , Sam Hashemi , Tobias Heimann , Fucang Jia , Jiexin Jiang , Emily Kaczmarek , Chris Kang , Seung Kwan Kang , Mohammad Khazaei , Julien Khlaut , Petros Koutsouvelis , Jae Sung Lee , Yuchong Li , Mengye Lyu , Mingchen Ma , Anant Madabhushi , Klaus H. Maier-Hein , Pierre Manceron , Andrés Martínez Mora , Moona Mazher , Felix Meister , Nataliia Molchanova , Steven A. Niederer , Leonard Nürnberg , Jinah Park , Abdul Qayyum , Jonas Richiardi , Antoine Saporta , Branislav Setlak , Ning Shen , Justin Szeto , Constantin Ulrich , Puru Vaish , Vibujithan Vigneshwaran , Leroy Volmer , Zihao Wang , Siqi Wei , Anthony Winder , Jelmer M. Wolterink , Maxence Wynen , Chang Yang , Si Young Yie , Mostafa Mehdipour Ghazi , Akshay Pai , Espen Jimenez Solem , Sebastian Nørgaard Llambias , Mikael Boesen , Michael Eriksen Benros , Juan Eugenio Iglesias , Mads Nielsen

As revealed by the scaling law of fine-grained MoE, model performance ceases to be improved once the granularity of the intermediate dimension exceeds the optimal threshold, limiting further gains from single-dimension fine-grained design.…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Ning Liao , Xiaoxing Wang , Xiaohan Qin , Junchi Yan

Video inpainting is the task of filling a region in a video in a visually convincing manner. It is very challenging due to the high dimensionality of the data and the temporal consistency required for obtaining convincing results. Recently,…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Nicolas Cherel , Andrés Almansa , Yann Gousseau , Alasdair Newson

Biometric capture devices have been utilised to estimate a person's alertness through near-infrared iris images, expanding their use beyond just biometric recognition. However, capturing a substantial number of corresponding images related…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Juan E. Tapia , Christoph Busch

Remote sensing images present unique challenges to image analysis due to the extensive geographic coverage, hardware limitations, and misaligned multi-scale images. This paper revisits the classical multi-scale representation learning…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Maofeng Tang , Andrei Cozma , Konstantinos Georgiou , Hairong Qi

Face recognition (FR) stands as one of the most crucial applications in computer vision. The accuracy of FR models has significantly improved in recent years due to the availability of large-scale human face datasets. However, directly…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Xiao Lin , Yuge Huang , Jianqing Xu , Yuxi Mi , Shuigeng Zhou , Shouhong Ding

Understanding whether self-supervised learning methods can scale with unlimited data is crucial for training large-scale models. In this work, we conduct an empirical study on the scaling capability of masked image modeling (MIM) methods…

计算机视觉与模式识别 · 计算机科学 2023-05-25 Cheng-Ze Lu , Xiaojie Jin , Qibin Hou , Jun Hao Liew , Ming-Ming Cheng , Jiashi Feng

In the past decade, image foundation models (IFMs) have achieved unprecedented progress. However, the potential of directly using IFMs for video self-supervised representation learning has largely been overlooked. In this study, we propose…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Jingwei Wu , Zhewei Huang , Chang Liu

Learning transferable representations from unlabeled time series is crucial for improving performance in data-scarce classification. Existing self-supervised methods often operate at the point level and rely on unidirectional encoding,…

机器学习 · 计算机科学 2026-03-02 Mingyue Cheng , Xiaoyu Tao , Zhiding Liu , Qi Liu , Hao Zhang , Rujiao Zhang , Enhong Chen

In this paper, we propose a Text-Degradation Invariant Auto Encoder (Text-DIAE), a self-supervised model designed to tackle two tasks, text recognition (handwritten or scene-text) and document image enhancement. We start by employing a…

Automated analysis of surgical videos is crucial for improving surgical training, workflow optimization, and postoperative assessment. We introduce a CSMAE, Masked Autoencoder (MAE)-based pretraining approach, specifically developed for…

计算机视觉与模式识别 · 计算机科学 2025-02-14 Nisarg A. Shah , Wele Gedara Chaminda Bandara , Shameema Skider , S. Swaroop Vedula , Vishal M. Patel

This paper introduces the Efficient Decoupled Masked Autoencoder (EDMAE), a novel self-supervised method for recognizing standard views in pediatric echocardiography. EDMAE introduces a new proxy task based on the encoder-decoder structure.…

图像与视频处理 · 电气工程与系统科学 2023-08-04 Yiman Liu , Xiaoxiang Han , Tongtong Liang , Bin Dong , Jiajun Yuan , Menghan Hu , Qiaohong Liu , Jiangang Chen , Qingli Li , Yuqi Zhang

We survey applications of pretrained foundation models in robotics. Traditional deep learning models in robotics are trained on small datasets tailored for specific tasks, which limits their adaptability across diverse applications. In…

Infrared and visible image fusion has emerged as a prominent research area in computer vision. However, little attention has been paid to the fusion task in complex scenes, leading to sub-optimal results under interference. To fill this…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Xilai Li , Xiaosong Li , Tianshu Tan , Huafeng Li , Tao Ye

We present an extension to masked autoencoders (MAE) which improves on the representations learnt by the model by explicitly encouraging the learning of higher scene-level features. We do this by: (i) the introduction of a perceptual…

计算机视觉与模式识别 · 计算机科学 2023-03-29 Samyakh Tukra , Frederick Hoffman , Ken Chatfield

Infrared image helps improve the perception capabilities of autonomous driving in complex weather conditions such as fog, rain, and low light. However, infrared image often suffers from low contrast, especially in non-heat-emitting targets…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Siyuan Chai , Xiaodong Guo , Tong Liu

Image inpainting is a technique used to restore missing or damaged regions of an image. Traditional methods primarily utilize information from adjacent pixels for reconstructing missing areas, while they struggle to preserve complex details…

计算机视觉与模式识别 · 计算机科学 2025-04-25 Junyan Zhang , Yan Li , Mengxiao Geng , Liu Shi , Qiegen Liu

Recent advancements in artificial intelligence (AI), particularly foundation models (FMs), have revolutionized medical image analysis, demonstrating strong zero- and few-shot performance across diverse medical imaging tasks, from…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Praveenbalaji Rajendran , Mojtaba Safari , Wenfeng He , Mingzhe Hu , Shansong Wang , Jun Zhou , Xiaofeng Yang