中文
相关论文

相关论文: Multi-modal, multi-scale representation learning f…

200 篇论文

Very Long Baseline Interferometry (VLBI) astrometry is a well established technique for achieving $\pm10~\mu$as parallax accuracies at frequencies well above 10~GHz. At lower frequencies, uncompensated interferometer delays associated with…

Positional encodings are a core part of transformer-based models, enabling processing of sequential data without recurrence. This paper presents a theoretical framework to analyze how various positional encoding methods, including…

机器学习 · 计算机科学 2025-06-10 Yin Li

The fusion of LiDARs and cameras has been increasingly adopted in autonomous driving for perception tasks. The performance of such fusion-based algorithms largely depends on the accuracy of sensor calibration, which is challenging due to…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Yuxuan Xiao , Yao Li , Chengzhen Meng , Xingchen Li , Jianmin Ji , Yanyong Zhang

Solar irradiance is fundamental data crucial for analyses related to weather and climate. High-precision estimation models are necessary to create areal data for solar irradiance. In this study, we developed a novel estimation model by…

大气与海洋物理 · 物理学 2024-07-08 Jun Sasaki , Maki Okada , Kenji Utsunomiya , Koji Yamaguchi

Bayesian spatial modeling provides a flexible framework for whole-brain fMRI analysis by explicitly incorporating spatial dependencies, overcoming the limitations of traditional massive univariate approaches that lead to information waste.…

统计方法学 · 统计学 2025-11-18 Yuan Zhong , Gang Chen , Paul A. Taylor , Jian Kang

This study presents a multisensory machine learning architecture for object recognition by employing a novel dataset that was constructed with the iCub robot, which is equipped with three cameras and a depth sensor. The proposed…

机器人学 · 计算机科学 2020-09-15 Murat Kirtay , Guido Schillaci , Verena V. Hafner

In the era of multinational cooperation, gathering and analyzing the satellite images are getting easier and more important. Typical procedure of the satellite image analysis include transmission of the bulky image data from satellite to…

图像与视频处理 · 电气工程与系统科学 2022-07-25 KyungChae Lee

Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such representations directly from unposed multi-view images…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Bo Zhou , Qiuxia Lai , Zeren Sun , Xiangbo Shu , Yazhou Yao , Wenguan Wang

The primary aim of this manuscript is to underscore a significant limitation in current deep learning models, particularly vision models. Unlike human vision, which efficiently selects only the essential visual areas for further processing,…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Ali Borji

The log-ratio (LR) operator has been widely employed to generate the difference image for synthetic aperture radar (SAR) image change detection. However, the difference image generated by this pixel-wise operator can be subject to SAR…

计算机视觉与模式识别 · 计算机科学 2020-02-19 Rongfang Wang , Jia-Wei Chen , Yule Wang , Licheng Jiao , Mi Wang

The next generation of Earth observation satellites will seek to deploy intelligent models directly onboard the payload in order to minimize the latency incurred by the transmission and processing chain of the ground segment, for…

图像与视频处理 · 电气工程与系统科学 2026-01-09 Ziyao Yi , Davide Piccinini , Diego Valsesia , Tiziano Bianchi , Enrico Magli

Hyperspectral image (HSI) classification presents inherent challenges due to high spectral dimensionality, significant domain shifts, and limited availability of labeled data. To address these issues, we propose a novel Active Transfer…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Muhammad Ahmad , Francesco Mauro , Manuel Mazzara , Salvatore Distefano , Adil Mehmood Khan , Silvia Liberata Ullo

Recent advancements in learned image compression (LIC) methods have demonstrated superior performance over traditional hand-crafted codecs. These learning-based methods often employ convolutional neural networks (CNNs) or Transformer-based…

计算机视觉与模式识别 · 计算机科学 2024-08-08 Hamidreza Soltani , Erfan Ghasemi

Compressive imaging is an emerging application of compressed sensing, devoted to acquisition, encoding and reconstruction of images using random projections as measurements. In this paper we propose a novel method to provide a scalable…

信息论 · 计算机科学 2013-10-07 Diego Valsesia , Enrico Magli

In this work we investigate the viability of foundational AI/ML models for Synthetic Aperture Radar (SAR) object recognition tasks. We are inspired by the tremendous progress being made in the wider community, particularly in the natural…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Nathan Inkawhich

Convolution is spatially-symmetric, i.e., the visual features are independent of its position in the image, which limits its ability to utilize contextual cues for visual recognition. This paper addresses this issue by introducing a…

计算机视觉与模式识别 · 计算机科学 2018-04-04 Yan Wang , Lingxi Xie , Siyuan Qiao , Ya Zhang , Wenjun Zhang , Alan L. Yuille

Modern microscopy routinely produces gigapixel images that contain structures across multiple spatial scales, from fine cellular morphology to broader tissue organization. Many analysis tasks require combining these scales, yet most vision…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Albert Dominguez Mantes , Gioele La Manno , Martin Weigert

Semantic segmentation of satellite imagery is crucial for Earth observation applications, but remains constrained by limited labelled training data. While self-supervised pretraining methods like Masked Autoencoders (MAE) have shown…

计算机视觉与模式识别 · 计算机科学 2025-07-17 John Waithaka , Moise Busogi

The field of computer vision is undergoing a paradigm shift toward large-scale foundation model pre-training via self-supervised learning (SSL). Leveraging large volumes of unlabeled brain MRI data, such models can learn anatomical priors…

图像与视频处理 · 电气工程与系统科学 2026-01-15 Petros Koutsouvelis , Matej Gazda , Leroy Volmer , Sina Amirrajab , Kamil Barbierik , Branislav Setlak , Jakub Gazda , Peter Drotar

Massive multiple-input multiple-output (MIMO) is a key enabler for the high data rates required by the sixth-generation networks, yet its performance hinges on effective beam management with low training overhead. This paper proposes an…

信息论 · 计算机科学 2026-02-27 Yijie Bian , Wei Guo , Jie Yang , Shenghui Song , Jun Zhang , Shi Jin , Khaled B. Letaief