English
Related papers

Related papers: Multi-modal, multi-scale representation learning f…

200 papers

Very Long Baseline Interferometry (VLBI) astrometry is a well established technique for achieving $\pm10~\mu$as parallax accuracies at frequencies well above 10~GHz. At lower frequencies, uncompensated interferometer delays associated with…

Instrumentation and Methods for Astrophysics · Physics 2023-02-15 Lucas J. Hyland , Mark J. Reid , Simon P. Ellingsen , Maria J. Rioja , Richard Dodson , Gabor Orosz , Colin R. Masson , Jamie M. McCallum

Positional encodings are a core part of transformer-based models, enabling processing of sequential data without recurrence. This paper presents a theoretical framework to analyze how various positional encoding methods, including…

Machine Learning · Computer Science 2025-06-10 Yin Li

The fusion of LiDARs and cameras has been increasingly adopted in autonomous driving for perception tasks. The performance of such fusion-based algorithms largely depends on the accuracy of sensor calibration, which is challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Yuxuan Xiao , Yao Li , Chengzhen Meng , Xingchen Li , Jianmin Ji , Yanyong Zhang

Solar irradiance is fundamental data crucial for analyses related to weather and climate. High-precision estimation models are necessary to create areal data for solar irradiance. In this study, we developed a novel estimation model by…

Atmospheric and Oceanic Physics · Physics 2024-07-08 Jun Sasaki , Maki Okada , Kenji Utsunomiya , Koji Yamaguchi

Bayesian spatial modeling provides a flexible framework for whole-brain fMRI analysis by explicitly incorporating spatial dependencies, overcoming the limitations of traditional massive univariate approaches that lead to information waste.…

Methodology · Statistics 2025-11-18 Yuan Zhong , Gang Chen , Paul A. Taylor , Jian Kang

This study presents a multisensory machine learning architecture for object recognition by employing a novel dataset that was constructed with the iCub robot, which is equipped with three cameras and a depth sensor. The proposed…

Robotics · Computer Science 2020-09-15 Murat Kirtay , Guido Schillaci , Verena V. Hafner

In the era of multinational cooperation, gathering and analyzing the satellite images are getting easier and more important. Typical procedure of the satellite image analysis include transmission of the bulky image data from satellite to…

Image and Video Processing · Electrical Eng. & Systems 2022-07-25 KyungChae Lee

Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such representations directly from unposed multi-view images…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Bo Zhou , Qiuxia Lai , Zeren Sun , Xiangbo Shu , Yazhou Yao , Wenguan Wang

The primary aim of this manuscript is to underscore a significant limitation in current deep learning models, particularly vision models. Unlike human vision, which efficiently selects only the essential visual areas for further processing,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Ali Borji

The log-ratio (LR) operator has been widely employed to generate the difference image for synthetic aperture radar (SAR) image change detection. However, the difference image generated by this pixel-wise operator can be subject to SAR…

Computer Vision and Pattern Recognition · Computer Science 2020-02-19 Rongfang Wang , Jia-Wei Chen , Yule Wang , Licheng Jiao , Mi Wang

The next generation of Earth observation satellites will seek to deploy intelligent models directly onboard the payload in order to minimize the latency incurred by the transmission and processing chain of the ground segment, for…

Image and Video Processing · Electrical Eng. & Systems 2026-01-09 Ziyao Yi , Davide Piccinini , Diego Valsesia , Tiziano Bianchi , Enrico Magli

Hyperspectral image (HSI) classification presents inherent challenges due to high spectral dimensionality, significant domain shifts, and limited availability of labeled data. To address these issues, we propose a novel Active Transfer…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Muhammad Ahmad , Francesco Mauro , Manuel Mazzara , Salvatore Distefano , Adil Mehmood Khan , Silvia Liberata Ullo

Recent advancements in learned image compression (LIC) methods have demonstrated superior performance over traditional hand-crafted codecs. These learning-based methods often employ convolutional neural networks (CNNs) or Transformer-based…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Hamidreza Soltani , Erfan Ghasemi

Compressive imaging is an emerging application of compressed sensing, devoted to acquisition, encoding and reconstruction of images using random projections as measurements. In this paper we propose a novel method to provide a scalable…

Information Theory · Computer Science 2013-10-07 Diego Valsesia , Enrico Magli

In this work we investigate the viability of foundational AI/ML models for Synthetic Aperture Radar (SAR) object recognition tasks. We are inspired by the tremendous progress being made in the wider community, particularly in the natural…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Nathan Inkawhich

Convolution is spatially-symmetric, i.e., the visual features are independent of its position in the image, which limits its ability to utilize contextual cues for visual recognition. This paper addresses this issue by introducing a…

Computer Vision and Pattern Recognition · Computer Science 2018-04-04 Yan Wang , Lingxi Xie , Siyuan Qiao , Ya Zhang , Wenjun Zhang , Alan L. Yuille

Modern microscopy routinely produces gigapixel images that contain structures across multiple spatial scales, from fine cellular morphology to broader tissue organization. Many analysis tasks require combining these scales, yet most vision…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Albert Dominguez Mantes , Gioele La Manno , Martin Weigert

Semantic segmentation of satellite imagery is crucial for Earth observation applications, but remains constrained by limited labelled training data. While self-supervised pretraining methods like Masked Autoencoders (MAE) have shown…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 John Waithaka , Moise Busogi

The field of computer vision is undergoing a paradigm shift toward large-scale foundation model pre-training via self-supervised learning (SSL). Leveraging large volumes of unlabeled brain MRI data, such models can learn anatomical priors…

Image and Video Processing · Electrical Eng. & Systems 2026-01-15 Petros Koutsouvelis , Matej Gazda , Leroy Volmer , Sina Amirrajab , Kamil Barbierik , Branislav Setlak , Jakub Gazda , Peter Drotar

Massive multiple-input multiple-output (MIMO) is a key enabler for the high data rates required by the sixth-generation networks, yet its performance hinges on effective beam management with low training overhead. This paper proposes an…

Information Theory · Computer Science 2026-02-27 Yijie Bian , Wei Guo , Jie Yang , Shenghui Song , Jun Zhang , Shi Jin , Khaled B. Letaief