中文
相关论文

相关论文: TerraMind: Large-Scale Generative Multimodality fo…

200 篇论文

Self-supervised pre-training bears potential to generate expressive representations without human annotation. Most pre-training in Earth observation (EO) are based on ImageNet or medium-size, labeled remote sensing (RS) datasets. We share…

计算机视觉与模式识别 · 计算机科学 2023-05-30 Yi Wang , Nassim Ait Ali Braham , Zhitong Xiong , Chenying Liu , Conrad M Albrecht , Xiao Xiang Zhu

Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images from various viewpoints, textual descriptions, geographic coordinates, etc.). Current…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Oskar Kristoffersen , Alba Reinders Sánchez , Morten Rieger Hannemose , Anders Bjorholm Dahl , Dim P. Papadopoulos

In the era of deep learning, annotated datasets have become a crucial asset to the remote sensing community. In the last decade, a plethora of different datasets was published, each designed for a specific data type and with a specific task…

计算机视觉与模式识别 · 计算机科学 2022-09-27 Michael Schmitt , Pedram Ghamisi , Naoto Yokoya , Ronny Hänsch

Integrated sensing and communications is a key enabler for the 6G wireless communication systems. The multiple sensing modalities will allow the base station to have a more accurate representation of the environment, leading to…

计算机视觉与模式识别 · 计算机科学 2024-06-28 Mohammad Farzanullah , Han Zhang , Akram Bin Sediq , Ali Afana , Melike Erol-Kantarci

Current Large Multimodal Models (LMMs) in Earth Observation typically neglect the critical "vertical" dimension, limiting their reasoning capabilities in complex remote sensing geometries and disaster scenarios where physical spatial…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Xuran Hu , Zhitong Xiong , Zhongcheng Hong , Yifang Ban , Xiaoxiang Zhu , Wufan Zhao

Multi-modality data is becoming readily available in remote sensing (RS) and can provide complementary information about the Earth's surface. Effective fusion of multi-modal information is thus important for various applications in RS, but…

计算机视觉与模式识别 · 计算机科学 2022-07-27 Qinghui Liu , Michael Kampffmeyer , Robert Jenssen , Arnt-Børre Salberg

The CREATE database is composed of 14 hours of multimodal recordings from a mobile robotic platform based on the iRobot Create. The various sensors cover vision, audition, motors and proprioception. The dataset has been designed in the…

机器人学 · 计算机科学 2018-02-01 Simon Brodeur , Simon Carrier , Jean Rouat

With the continuous improvement of computing power and deep learning algorithms in recent years, the foundation model has grown in popularity. Because of its powerful capabilities and excellent performance, this technology is being adopted…

计算机视觉与模式识别 · 计算机科学 2023-06-08 Yifeng Shi , Feng Lv , Xinliang Wang , Chunlong Xia , Shaojie Li , Shujie Yang , Teng Xi , Gang Zhang

Accurate semantic segmentation of remote sensing imagery is critical for various Earth observation applications, such as land cover mapping, urban planning, and environmental monitoring. However, individual data sources often present…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Ivica Dimitrovski , Vlatko Spasev , Ivan Kitanovski

We introduce EarthPT -- an Earth Observation (EO) pretrained transformer. EarthPT is a 700 million parameter decoding transformer foundation model trained in an autoregressive self-supervised manner and developed specifically with EO…

机器学习 · 计算机科学 2024-01-12 Michael J. Smith , Luke Fleming , James E. Geach

Jointly harnessing complementary features of multi-modal input data in a common latent space has been found to be beneficial long ago. However, the influence of each modality on the models decision remains a puzzle. This study proposes a…

计算机视觉与模式识别 · 计算机科学 2023-04-06 Burak Ekim , Michael Schmitt

Many adaptations of transformers have emerged to address the single-modal vision tasks, where self-attention modules are stacked to handle input sources like images. Intuitively, feeding multiple modalities of data to vision transformers…

计算机视觉与模式识别 · 计算机科学 2022-07-18 Yikai Wang , Xinghao Chen , Lele Cao , Wenbing Huang , Fuchun Sun , Yunhe Wang

As repositories of large scale data in earth observation (EO) have grown, so have transfer and storage costs for model training and inference, expending significant resources. We introduce Neural Embedding Compression (NEC), based on the…

机器学习 · 计算机科学 2024-07-11 Carlos Gomes , Thomas Brunschwiler

Semantic segmentation of multi-modal remote sensing imagery plays a pivotal role in land use/land cover (LULC) mapping, environmental monitoring, and precision earth observation. Current multi-modal approaches mainly focus on integrating…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Jinkun Dai , Yuanxin Ye , Peng Tang , Tengfeng Tang , Xianping Ma , Jing Xiao , Mi Wang

With the ever-increasing volumes of the Earth observation data present in the archives of large programmes such as Copernicus, there is a growing need for efficient vector representations of the underlying raw data. The approach of…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Mikolaj Czerkawski , Marcin Kluczek , Jędrzej S. Bojanowski

Pre-trained Foundation Models (PFMs) have ushered in a paradigm-shift in Artificial Intelligence, due to their ability to learn general-purpose representations that can be readily employed in a wide range of downstream tasks. While PFMs…

数据库 · 计算机科学 2024-11-13 Pasquale Balsebre , Weiming Huang , Gao Cong , Yi Li

Animal re-identification (ReID) faces critical challenges due to viewpoint variations, particularly in Aerial-Ground (AG-ReID) settings where models must match individuals across drastic elevation changes. However, existing datasets lack…

计算机视觉与模式识别 · 计算机科学 2026-05-29 William Grolleau , Achraf Chaouch , Astrid Sabourin , Guillaume Lapouge , Catherine Achard

This project performs multimodal sentiment analysis using the CMU-MOSEI dataset, using transformer-based models with early fusion to integrate text, audio, and visual modalities. We employ BERT-based encoders for each modality, extracting…

计算与语言 · 计算机科学 2025-07-16 Jugal Gajjar , Kaustik Ranaware

The fast simulation of dynamical systems is a key challenge in many scientific and engineering applications, such as weather forecasting, disease control, and drug discovery. With the recent success of deep learning, there is increasing…

机器学习 · 计算机科学 2024-10-02 Zezheng Song , Jiaxin Yuan , Haizhao Yang

Recent progress in self-supervision shows that pre-training large neural networks on vast amounts of unsupervised data can lead to impressive increases in generalisation for downstream tasks. Such models, recently coined as foundation…