English
Related papers

Related papers: Multi-modal, multi-scale representation learning f…

200 papers

Deep learning models have achieved excellent recognition results on large-scale video benchmarks. However, they perform poorly when applied to videos with rare scenes or objects, primarily due to the bias of existing video datasets. We…

Computer Vision and Pattern Recognition · Computer Science 2022-09-21 Haodong Duan , Yue Zhao , Kai Chen , Yuanjun Xiong , Dahua Lin

Recent advancements in foundation models, typically trained with self-supervised learning on large-scale and diverse datasets, have shown great potential in medical image analysis. However, due to the significant spatial heterogeneity of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-25 Lingxiao Luo , Xuanzhong Chen , Bingda Tang , Xinsheng Chen , Rong Han , Chengpeng Hu , Yujiang Li , Ting Chen

The rapid advancement of autonomous systems, including self-driving vehicles and drones, has intensified the need to forge true Spatial Intelligence from multi-modal onboard sensor data. While foundation models excel in single-modal…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Song Wang , Lingdong Kong , Xiaolu Liu , Hao Shi , Wentong Li , Jianke Zhu , Steven C. H. Hoi

Super-resolution is aimed at reconstructing high-resolution images from low-resolution observations. State-of-the-art approaches underpinned with deep learning allow for obtaining outstanding results, generating images of high perceptual…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Maciej Ziaja , Pawel Kowaleczko , Daniel Kostrzewa , Nicolas Longépé , Michal Kawulok

Foundation models are vital tools in various Computer Vision applications. They take as input a single RGB image and output a deep feature representation that is useful for various applications. However, in case we have multiple views of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Leo Segre , Or Hirschorn , Shai Avidan

A number of applications, such as mobile robots or automated vehicles, use LiDAR sensors to obtain detailed information about their three-dimensional surroundings. Many methods use image-like projections to efficiently process these LiDAR…

Computer Vision and Pattern Recognition · Computer Science 2021-12-01 Larissa T. Triess , David Peter , J. Marius Zöllner

As the important component of the Earth observation system, hyperspectral imaging satellites provide high-fidelity and enriched information for the formulation of related policies due to the powerful spectral measurement capabilities.…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Li Fang , Tianyu Li , Yanghong Lin , Shudong Zhou , Wei Yao

Direct Low Earth Orbit satellite-to-handheld links are expected to be part of a new era in satellite communications. Space-Division Multiple Access precoding is a technique that reduces interference among satellite beams, therefore…

Signal Processing · Electrical Eng. & Systems 2023-03-22 Steffen Gracla , Alea Schröder , Maik Röper , Carsten Bockelmann , Dirk Wübben , Armin Dekorsy

We introduce Multi-scale Adaptive Recurrent Biomedical Linear-time Encoder (MARBLE), the first \textit{purely Mamba-based} multi-state multiple instance learning (MIL) framework for whole-slide image (WSI) analysis. MARBLE processes…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Jagan Mohan Reddy Dwarampudi , Joshua Wong , Hien Van Nguyen , Tania Banerjee

Bias in machine learning models can lead to unfair decision making, and while it has been well-studied in the image and text domains, it remains underexplored in action recognition. Action recognition models often suffer from background…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Joseph Fioresi , Ishan Rajendrakumar Dave , Mubarak Shah

Surveillance systems play a critical role in security and reconnaissance, but their performance is often compromised by low-quality images and videos, leading to reduced accuracy in face recognition. Additionally, existing AI-based facial…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Anees Nashath Shaik , Barbara Villarini , Vasileios Argyriou

This work presents a systematic investigation of custom convolutional neural network architectures for satellite land use classification, achieving 97.23% test accuracy on the EuroSAT dataset without reliance on pre-trained models. Through…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Aditya Vir

Vision foundation models in remote sensing have been extensively studied due to their superior generalization on various downstream tasks. Synthetic Aperture Radar (SAR) offers all-day, all-weather imaging capabilities, providing…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Mengyu Wang , Hanbo Bi , Yingchao Feng , Linlin Xin , Shuo Gong , Tianqi Wang , Zhiyuan Yan , Peijin Wang , Wenhui Diao , Xian Sun

In this article, the analysis of existing models of satellite image recognition was carried out, the problems in the field of satellite image recognition as a source of information were considered and analyzed, deep learning methods were…

Computer Vision and Pattern Recognition · Computer Science 2022-12-08 Alexey Averkin , Sergey Yarushev

We address the issue of visual saliency from three perspectives. First, we consider saliency detection as a frequency domain analysis problem. Second, we achieve this by employing the concept of {\it non-saliency}. Third, we simultaneously…

Computer Vision and Pattern Recognition · Computer Science 2016-05-09 Jian Li , Martin Levine , Xiangjing An , Xin Xu , Hangen He

Sampling low-energy molecular conformations, spatial arrangements of atoms in a molecule, is a critical task for many different calculations performed in the drug discovery and optimization process. Numerous specialized equivariant networks…

Biomolecules · Quantitative Biology 2025-06-25 Viatcheslav Gurev , Timothy Rumbell

Combining satellite imagery with machine learning (SIML) has the potential to address global challenges by remotely estimating socioeconomic and environmental conditions in data-poor regions, yet the resource requirements of SIML limit its…

Recently, the philosophy of visual saliency and attention has started to gain popularity in the robotics community. Therefore, this paper aims to mimic this mechanism in SLAM framework by using saliency prediction model. Comparing with…

Robotics · Computer Science 2020-12-23 Ke Wang , Sai Ma , Junlan Chen , Jianbo Lu

We present a novel framework using Energy-Based Models (EBMs) for localizing a ground vehicle mounted with a range sensor against satellite imagery in the absence of GPS. Lidar sensors have become ubiquitous on autonomous vehicles for…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Alan Wu , Michael S. Ryoo

In this paper, we address the challenge of image resolution variation for the Segment Anything Model (SAM). SAM, known for its zero-shot generalizability, exhibits a performance degradation when faced with datasets with varying image sizes.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Yiran Song , Qianyu Zhou , Xiangtai Li , Deng-Ping Fan , Xuequan Lu , Lizhuang Ma