中文
相关论文

相关论文: GeoLink: Empowering Remote Sensing Foundation Mode…

200 篇论文

Geometry problem-solving (GPS), a challenging task requiring both visual comprehension and symbolic reasoning, effectively measures the reasoning capabilities of multimodal large language models (MLLMs). Humans exhibit strong reasoning…

计算与语言 · 计算机科学 2025-04-25 Liangyu Xu , Yingxiu Zhao , Jingyun Wang , Yingyao Wang , Bu Pi , Chen Wang , Mingliang Zhang , Jihao Gu , Xiang Li , Xiaoyong Zhu , Jun Song , Bo Zheng

Similar to vision-and-language navigation (VLN) tasks that focus on bridging the gap between vision and language for embodied navigation, the new Rendezvous (RVS) task requires reasoning over allocentric spatial relationships (independent…

计算与语言 · 计算机科学 2024-07-01 Tzuf Paz-Argaman , John Palowitch , Sayali Kulkarni , Reut Tsarfaty , Jason Baldridge

Multi-spectral imagery is a valuable input signal for Remote Sensing applications, such as land-use and land-cover classification and environmental monitoring. However, generalist Large Multi-modal Models (LMMs) are typically trained on RGB…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Dahun Kim , Ganesh Satish Mallya , Anelia Angelova

Road network data provides rich information about cities, but processing worldwide OpenStreetMap (OSM) data is computationally intensive, and the resulting graphs are often difficult to unify for benchmarking downstream tasks. Existing…

数据库 · 计算机科学 2026-05-22 Guanjie Zheng , Ziyang Su , Yiheng Wang , Yuhang Luo , Hongwei Zhang , Xuanhe Zhou , Linghe Kong , Fan Wu , Wen Ling

Traditional robot navigation systems primarily utilize occupancy grid maps and laser-based sensing technologies, as demonstrated by the popular move_base package in ROS. Unlike robots, humans navigate not only through spatial awareness and…

机器人学 · 计算机科学 2025-07-22 Fujing Xie , Jiajie Zhang , Sören Schwertfeger

Deep Learning (DL) is undergoing a paradigm shift with the emergence of foundation models. In this work, we focus on Contrastive Language-Image Pre-training (CLIP), a Vision-Language foundation model that achieves high accuracy across…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Angelos Zavras , Dimitrios Michail , Begüm Demir , Ioannis Papoutsis

Simultaneous Localization and Mapping (SLAM) is a foundational component in robotics, AR/VR, and autonomous systems. With the rising focus on spatial AI in recent years, combining SLAM with semantic understanding has become increasingly…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Jisang Yoo , Gyeongjin Kang , Hyun-kyu Ko , Hyeonwoo Yu , Eunbyung Park

Training specific deep learning models for particular tasks is common across various domains within seismology. However, this approach encounters two limitations: inadequate labeled data for certain tasks and limited generalization across…

地球物理 · 物理学 2023-09-06 Xu Si , Xinming Wu , Hanlin Sheng , Jun Zhu , Zefeng Li

Remote sensing (RS) cross-modal text-image retrieval has attracted extensive attention for its advantages of flexible input and efficient query. However, traditional methods ignore the characteristics of multi-scale and redundant targets in…

计算机视觉与模式识别 · 计算机科学 2022-04-22 Zhiqiang Yuan , Wenkai Zhang , Kun Fu , Xuan Li , Chubo Deng , Hongqi Wang , Xian Sun

Multimodal change detection (MMCD) identifies changed areas in multimodal remote sensing (RS) data, demonstrating significant application value in land use monitoring, disaster assessment, and urban sustainable development. However,…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Xuanguang Liu , Lei Ding , Yujie Li , Chenguang Dai , Zhenchao Zhang , Mengmeng Li , Ziyi Yang , Yifan Sun , Yongqi Sun , Hanyun Wang

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks. However, improving their spatial intelligence remains a challenge. Existing 3D MLLMs always rely on additional 3D or…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Diankun Wu , Fangfu Liu , Yi-Hsin Hung , Yueqi Duan

The "thinking-with-images" paradigm enables multimodal large language models (MLLMs) to actively explore visual scenes via zoom-in tools. This is essential for ultra-high-resolution (UHR) remote sensing VQA, where task-relevant cues are…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Fengxiang Wang , Mingshuo Chen , Yueying Li , Yajie Yang , Yifan Zhang , Long Lan , Xue Yang , Hongda Sun , Yulin Wang , Di Wang , Jun Song , Jing Zhang , Bo Du

Artificial intelligence (AI) has significantly advanced Earth sciences, yet its full potential in to comprehensively modeling Earth's complex dynamics remains unrealized. Geoscience foundation models (GFMs) emerge as a paradigm-shifting…

人工智能 · 计算机科学 2024-11-13 Hao Zhang , Jin-Jian Xu , Hong-Wei Cui , Lin Li , Yaowen Yang , Chao-Sheng Tang , Niklas Boers

Vision foundation models have attracted significant attention for their ability to leverage large-scale unlabeled visual data. This advantage is particularly important in remote sensing, where data acquisition is costly and annotation often…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Hyobin Park , Minseok Seo , Dong-Geol Choi

Recent advances in foundation models have shown great promise in domains such as natural language processing and computer vision, and similar efforts are now emerging in the Earth Observation community. These models aim to generalize across…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Pierre Adorni , Minh-Tan Pham , Stéphane May , Sébastien Lefèvre

Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks but their remote sensing (RS) counterpart are relatively under explored. Unlike natural images, RS imagery presents unique challenges that…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Abduljaleel Adejumo , Faegheh Yeganli , Clifford Broni-bediako , Aoran Xiao , Naoto Yokoya , Mennatullah Siam

Earth Observation Foundation Models (EOFMs) have exploded in prevalence as tools for processing the massive volumes of remotely sensed and other earth observation data, and for delivering impact on the many essential earth monitoring tasks.…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Ryan P. Demilt , Nicholas LaHaye , Karis Tenneson

Recently, there has been increasing interest in multimodal applications that integrate text with other modalities, such as images, audio and video, to facilitate natural language interactions with multimodal AI systems. While applications…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Roger Ferrod , Luigi Di Caro , Dino Ienco

The mainstream paradigm of remote sensing image interpretation has long been dominated by vision-centered models, which rely on visual features for semantic understanding. However, these models face inherent limitations in handling…

人工智能 · 计算机科学 2026-01-28 Haifeng Li , Wang Guo , Haiyang Wu , Mengwei Wu , Jipeng Zhang , Qing Zhu , Yu Liu , Xin Huang , Chao Tao

Foundation models have reshaped the landscape of Remote Sensing (RS) by enhancing various image interpretation tasks. Pretraining is an active research topic, encompassing supervised and self-supervised learning methods to initialize model…

计算机视觉与模式识别 · 计算机科学 2024-05-31 Di Wang , Jing Zhang , Minqiang Xu , Lin Liu , Dongsheng Wang , Erzhong Gao , Chengxi Han , Haonan Guo , Bo Du , Dacheng Tao , Liangpei Zhang