中文
相关论文

相关论文: GeoDecoder: Empowering Multimodal Map Understandin…

200 篇论文

LLMs excel at linguistic tasks but lack the inner geospatial capabilities needed for time-critical disaster response, where reasoning about road networks, coordinates, and access to essential infrastructure such as hospitals, shelters, and…

计算与语言 · 计算机科学 2026-03-27 Ahmed El Fekih Zguir , Ferda Ofli , Muhammad Imran

We introduce GeoDiT, a diffusion transformer designed for text-to-satellite image generation with point-based control. Existing controlled satellite image generative models often require pixel-level maps that are time-consuming to acquire,…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Srikumar Sastry , Dan Cher , Brian Wei , Aayush Dhakal , Subash Khanal , Dev Gupta , Nathan Jacobs

Automated textual description of remote sensing images is crucial for unlocking their full potential in diverse applications, from environmental monitoring to urban planning and disaster management. However, existing studies in remote…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Kaiyu Li , Zixuan Jiang , Xiangyong Cao , Jiayu Wang , Yuchen Xiao , Deyu Meng , Zhi Wang

The increasing availability of geospatial foundation models has the potential to transform remote sensing applications such as land cover classification, environmental monitoring, and change detection. Despite promising benchmark results,…

Recent advancements in world models have revolutionized dynamic environment simulation, allowing systems to foresee future states and assess potential actions. In autonomous driving, these capabilities help vehicles anticipate the behavior…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Anthony Chen , Wenzhao Zheng , Yida Wang , Xueyang Zhang , Kun Zhan , Peng Jia , Kurt Keutzer , Shanghang Zhang

Ground Terrain Recognition is a difficult task as the context information varies significantly over the regions of a ground terrain image. In this paper, we propose a novel approach towards ground-terrain recognition via modeling the…

计算机视觉与模式识别 · 计算机科学 2020-10-28 Shuvozit Ghose , Pinaki Nath Chowdhury , Partha Pratim Roy , Umapada Pal

Historical geologic maps contain rich geospatial information, such as rock units, faults, folds, and bedding planes, that is critical for assessing mineral resources essential to renewable energy, electric vehicles, and national security.…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Weiwei Duan , Michael P. Gerlek , Steven N. Minton , Craig A. Knoblock , Fandel Lin , Theresa Chen , Leeje Jang , Sofia Kirsanova , Zekun Li , Yijun Lin , Yao-Yi Chiang

Sensemaking using automatically extracted information from text is a challenging problem. In this paper, we address a specific type of information extraction, namely extracting information related to descriptions of movement. Aggregating…

人机交互 · 计算机科学 2022-04-21 Scott Pezanowski , Prasenjit Mitra , Alan M. MacEachren

This work conducts an evaluation of GPT-4V's multimodal capability for medical image analysis, with a focus on three representative tasks of radiology report generation, medical visual question answering, and medical visual grounding. For…

计算机视觉与模式识别 · 计算机科学 2024-02-01 Yingshu Li , Yunyi Liu , Zhanyu Wang , Xinyu Liang , Lei Wang , Lingqiao Liu , Leyang Cui , Zhaopeng Tu , Longyue Wang , Luping Zhou

Graph-structured information offers rich contextual information that can enhance language models by providing structured relationships and hierarchies, leading to more expressive embeddings for various applications such as retrieval,…

Worldwide geo-localization involves determining the exact geographic location of images captured globally, typically guided by geographic cues such as climate, landmarks, and architectural styles. Despite advancements in geo-localization…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Furong Jia , Lanxin Liu , Ce Hou , Fan Zhang , Xinyan Liu , Yu Liu

Latent space geometry provides a rigorous and empirically valuable framework for interacting with the latent variables of deep generative models. This approach reinterprets Euclidean latent spaces as Riemannian through a pull-back metric,…

机器学习 · 统计学 2024-08-15 Stas Syrota , Pablo Moreno-Muñoz , Søren Hauberg

Mapping with uncertainty representation is required in many research domains, especially for localization. Although there are many investigations regarding the uncertainty of the pose estimation of an ego-robot with map information, the…

机器人学 · 计算机科学 2023-08-30 Qianqian Zou , Claus Brenner , Monika Sester

Pre-trained Foundation Models (PFMs) have ushered in a paradigm-shift in Artificial Intelligence, due to their ability to learn general-purpose representations that can be readily employed in a wide range of downstream tasks. While PFMs…

数据库 · 计算机科学 2024-11-13 Pasquale Balsebre , Weiming Huang , Gao Cong , Yi Li

GeoAI is evolving rapidly, fueled by diverse geospatial datasets like traffic patterns, environmental data, and crowdsourced OpenStreetMap (OSM) information. While sophisticated AI models are being developed, existing benchmarks are often…

We introduce a deep multitask architecture to integrate multityped representations of multimodal objects. This multitype exposition is less abstract than the multimodal characterization, but more machine-friendly, and thus is more precise…

机器学习 · 统计学 2016-03-07 Truyen Tran , Dinh Phung , Svetha Venkatesh

Single encoder-decoder methodologies for semantic segmentation are reaching their peak in terms of segmentation quality and efficiency per number of layers. To address these limitations, we propose a new architecture based on a decoder…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Gabriel L. Oliveira , Senthil Yogamani , Wolfram Burgard , Thomas Brox

In the area of geographic information processing. There are few researches on geographic text classification. However, the application of this task in Chinese is relatively rare. In our work, we intend to implement a method to extract text…

计算与语言 · 计算机科学 2021-01-28 Weipeng Jing , Xianyang Song , Donglin Di , Houbing Song

Multimodal large language models (MLLMs) have made significant progress in integrating visual and linguistic understanding. Existing benchmarks typically focus on high-level semantic capabilities, such as scene understanding and visual…

计算与语言 · 计算机科学 2025-02-18 Shangyu Xing , Changhao Xiang , Yuteng Han , Yifan Yue , Zhen Wu , Xinyu Liu , Zhangtai Wu , Fei Zhao , Xinyu Dai

Visual document understanding (VDU) has rapidly advanced with the development of powerful multi-modal language models. However, these models typically require extensive document pre-training data to learn intermediate representations and…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Souhail Bakkali , Sanket Biswas , Zuheng Ming , Mickaël Coustaty , Marçal Rusiñol , Oriol Ramos Terrades , Josep Lladós