English
Related papers

Related papers: Self-Supervised Multi-Modal World Model with 4D Sp…

200 papers

Significant efforts have been directed towards adapting self-supervised multimodal learning for Earth observation applications. However, most current methods produce coarse patch-sized embeddings, limiting their effectiveness and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Ibrahim Fayad , Max Zimmer , Martin Schwartz , Fabian Gieseke , Philippe Ciais , Gabriel Belouze , Sarah Brood , Aurelien De Truchis , Alexandre d'Aspremont

Landuse characterization is important for urban planning. It is traditionally performed with field surveys or manual photo interpretation, two practices that are time-consuming and labor-intensive. Therefore, we aim to automate landuse…

Computer Vision and Pattern Recognition · Computer Science 2019-05-07 Shivangi Srivastava , John E. Vargas-Muñoz , Devis Tuia

We introduce DeepCell, a novel circuit representation learning framework that effectively integrates multiview information from both And-Inverter Graphs (AIGs) and Post-Mapping (PM) netlists. At its core, DeepCell employs a self-supervised…

Machine Learning · Computer Science 2025-07-09 Zhengyuan Shi , Chengyu Ma , Ziyang Zheng , Lingfeng Zhou , Hongyang Pan , Wentao Jiang , Fan Yang , Xiaoyan Yang , Zhufei Chu , Qiang Xu

The shape of objects is an important source of visual information in a wide range of applications. One of the core challenges of shape quantification is to ensure that the extracted measurements remain invariant to transformations that…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Anna Foix Romero , Craig Russell , Alexander Krull , Virginie Uhlmann

Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x2013}$ action classification, ImageNet classification, etc. In…

Hyperspectral imaging provides precise classification for land use and cover due to its exceptional spectral resolution. However, the challenges of high dimensionality and limited spatial resolution hinder its effectiveness. This study…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Shivam Pande

Deep learning models are increasingly data-hungry, requiring significant resources to collect and compile the datasets needed to train them, with Earth Observation (EO) models being no exception. However, the landscape of datasets in EO is…

Computer Vision and Pattern Recognition · Computer Science 2024-06-24 Alistair Francis , Mikolaj Czerkawski

Masked Image Modeling has been one of the most popular self-supervised learning paradigms to learn representations from large-scale, unlabeled Earth Observation images. While incorporating multi-modal and multi-temporal Earth Observation…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Liang Zeng , Valerio Marsocci , Wufan Zhao , Andrea Nascetti , Maarten Vergauwen

We present convolutional neural network (CNN) based approaches for unsupervised multimodal subspace clustering. The proposed framework consists of three main stages - multimodal encoder, self-expressive layer, and multimodal decoder. The…

Machine Learning · Computer Science 2025-10-13 Mahdi Abavisani , Vishal M. Patel

In this paper, we propose a novel multimodal deep hashing neural decoder (MDHND) architecture, which integrates a deep hashing framework with a neural network decoder (NND) to create an effective multibiometric authentication system. The…

Computer Vision and Pattern Recognition · Computer Science 2019-03-08 Veeru Talreja , Sobhan Soleymani , Matthew C. Valenti , Nasser M. Nasrabadi

Attentional mechanisms are order-invariant. Positional encoding is a crucial component to allow attention-based deep model architectures such as Transformer to address sequences or images where the position of information matters. In this…

Machine Learning · Computer Science 2021-11-10 Yang Li , Si Si , Gang Li , Cho-Jui Hsieh , Samy Bengio

Dense prediction infers per-pixel values from a single image and is fundamental to 3D perception and robotics. Although real-world scenes exhibit strong structure, existing methods treat it as an independent pixel-wise prediction, often…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Seung Hyun Lee , Sangwoo Mo , Stella X. Yu

Neural fields excel in computer vision and robotics due to their ability to understand the 3D visual world such as inferring semantics, geometry, and dynamics. Given the capabilities of neural fields in densely representing a 3D scene from…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Muhammad Zubair Irshad , Sergey Zakharov , Vitor Guizilini , Adrien Gaidon , Zsolt Kira , Rares Ambrus

This work investigates the use of deep fully convolutional neural networks (DFCNN) for pixel-wise scene labeling of Earth Observation images. Especially, we train a variant of the SegNet architecture on remote sensing data over an urban…

Computer Vision and Pattern Recognition · Computer Science 2016-09-23 Nicolas Audebert , Bertrand Le Saux , Sébastien Lefèvre

Depth estimation is a crucial technology in robotics. Recently, self-supervised depth estimation methods have demonstrated great potential as they can efficiently leverage large amounts of unlabelled real-world data. However, most existing…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Siyu Chen , Hong Liu , Wenhao Li , Ying Zhu , Guoquan Wang , Jianbing Wu

Accurate four-dimensional (4D) precipitation information is essential for understanding the Earth's energy and water cycles, yet remains observationally unresolved at global scales. Conventional theory holds that geostationary infrared…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Tianchi Xu , Ziqiang Ma , Andrea Marinoni , Yuanpeng He , Xiaoqing Li , Chuanfeng Zhao , Kang He , Jintao Xu , Bohan Zhou , Wenbo Zhao , Haoshuang Chen , Tun Wang , Dongdong Wang , Yang Hong

Currently, this paper is under review in IEEE. Transformers have intrigued the vision research community with their state-of-the-art performance in natural language processing. With their superior performance, transformers have found their…

Computer Vision and Pattern Recognition · Computer Science 2022-10-05 Preetam Ghosh , Swalpa Kumar Roy , Bikram Koirala , Behnood Rasti , Paul Scheunders

Cross-view object geo-localization enables high-precision object localization through cross-view matching, with critical applications in autonomous driving, urban management, and disaster response. However, existing methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Shuhan Hu , Yiru Li , Yuanyuan Li , Yingying Zhu

Phase-field modeling is an effective but computationally expensive method for capturing the mesoscale morphological and microstructure evolution in materials. Hence, fast and generalizable surrogate models are needed to alleviate the cost…

Materials Science · Physics 2022-07-01 Vivek Oommen , Khemraj Shukla , Somdatta Goswami , Remi Dingreville , George Em Karniadakis

Recent progress in 4D implicit representation focuses on globally controlling the shape and motion with low dimensional latent vectors, which is prone to missing surface details and accumulating tracking error. While many deep local…

Computer Vision and Pattern Recognition · Computer Science 2022-08-19 Boyan Jiang , Xinlin Ren , Mingsong Dou , Xiangyang Xue , Yanwei Fu , Yinda Zhang
‹ Prev 1 3 4 5 6 7 10 Next ›