中文
相关论文

相关论文: GeoViSTA: Geospatial Vision-Tabular Transformer fo…

200 篇论文

Visual transformers have driven major progress in remote sensing image analysis, particularly in object detection and segmentation. Recent vision-language and multimodal models further extend these capabilities by incorporating auxiliary…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Yu Li , Guilherme N. DeSouza , Praveen Rao , Chi-Ren Shyu

Interpreting ultra-high-resolution (UHR) remote sensing images requires models to search for sparse and tiny visual evidence across large-scale scenes. Existing remote sensing vision-language models can inspect local regions with zooming…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Jiashun Zhu , Ronghao Fu , Jiasen Hu , Nachuan Xing , Xu Na , Xiao Yang , Zhiwen Lin , Weipeng Zhang , Lang Sun , Zhiheng Xue , Haoran Liu , Weijie Zhang , Bo Yang

There exists a correlation between geospatial activity temporal patterns and type of land use. A novel self-supervised approach is proposed to stratify landscape based on mobility activity time series. First, the time series signal is…

计算机视觉与模式识别 · 计算机科学 2024-01-18 Yi Cao , Swetava Ganguli , Vipul Pandey

Foundation models have transformed natural language processing and computer vision, and their impact is now reshaping remote sensing image analysis. With powerful generalization and transfer learning capabilities, they align naturally with…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Liling Yang , Ning Chen , Jun Yue , Yidan Liu , Jiayi Ma , Pedram Ghamisi , Antonio Plaza , Leyuan Fang

Recent advancements of image captioning have featured Visual-Semantic Fusion or Geometry-Aid attention refinement. However, those fusion-based models, they are still criticized for the lack of geometry information for inter and intra…

计算机视觉与模式识别 · 计算机科学 2021-09-30 Ling Cheng , Wei Wei , Feida Zhu , Yong Liu , Chunyan Miao

Existing methods for self-supervised representation learning of geospatial regions and map entities rely extensively on the design of pretext tasks, often involving augmentations or heuristic sampling of positive and negative pairs based on…

机器学习 · 计算机科学 2025-03-11 Theodor Lundqvist , Ludvig Delvret

Recent advances in multimodal large language models(MLLMs) have led to remarkable progress in visual grounding, enabling fine-grained cross-modal alignment between textual queries and image regions. However, transferring such capabilities…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Peirong Zhang , Yidan Zhang , Luxiao Xu , Jinliang Lin , Zonghao Guo , Fengxiang Wang , Xue Yang , Kaiwen Wei , Lei Wang

Managing natural resources and mitigating risks from floods, droughts, wildfires, and landslides require models that can accurately predict climate-driven land-surface responses. Traditional models often struggle with spatial generalization…

机器学习 · 计算机科学 2026-02-03 Nicholas Kraabel , Jiangtao Liu , Yuchen Bian , Daniel Kifer , Chaopeng Shen

We introduce TableVista, a comprehensive benchmark for evaluating foundation models in multimodal table reasoning under visual and structural complexity. TableVista consists of 3,000 high-quality table reasoning problems, where each…

计算与语言 · 计算机科学 2026-05-08 Zheyuan Yang , Liqiang Shang , Junjie Chen , Xun Yang , Chenglong Xu , Bo Yuan , Chenyuan Jiao , Yaoru Sun , Yilun Zhao

Accurate and robust navigation in unstructured environments requires fusing data from multiple sensors. Such fusion ensures that the robot is better aware of its surroundings, including areas of the environment that are not immediately…

机器人学 · 计算机科学 2024-03-12 Mateus Valverde Gasparino , Arun Narenthiran Sivakumar , Girish Chowdhary

Scalable general-purpose representations of the built environment are crucial for geospatial artificial intelligence applications. This paper introduces S2Vec, a novel self-supervised framework for learning such geospatial embeddings. S2Vec…

社会与信息网络 · 计算机科学 2026-01-08 Shushman Choudhury , Elad Aharoni , Chandrakumari Suvarna , Iveel Tsogsuren , Abdul Rahman Kreidieh , Chun-Ta Lu , Neha Arora

Robot localization remains a challenging task in GPS denied environments. State estimation approaches based on local sensors, e.g. cameras or IMUs, are drifting-prone for long-range missions as error accumulates. In this study, we aim to…

计算机视觉与模式识别 · 计算机科学 2022-05-17 Tianyi Zhang , Matthew Johnson-Roberson

Tables are widely used with various structures to organize and present data. Recent attempts on table understanding mainly focus on relational tables, yet overlook to other common table structures. In this paper, we propose TUTA, a unified…

信息检索 · 计算机科学 2021-07-21 Zhiruo Wang , Haoyu Dong , Ran Jia , Jia Li , Zhiyi Fu , Shi Han , Dongmei Zhang

Global localization is critical for autonomous navigation, particularly in scenarios where an agent must localize within a map generated in a different session or by another agent, as agents often have no prior knowledge about the…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Hannah Shafferman , Annika Thomas , Jouko Kinnari , Michael Ricard , Jose Nino , Jonathan How

We present GeoGrid-Bench, a benchmark designed to evaluate the ability of foundation models to understand geo-spatial data in the grid structure. Geo-spatial datasets pose distinct challenges due to their dense numerical values, strong…

Self-supervised learning has been shown to be very effective in learning useful representations, and yet much of the success is achieved in data types such as images, audio, and text. The success is mainly enabled by taking advantage of…

机器学习 · 计算机科学 2021-10-28 Talip Ucar , Ehsan Hajiramezanali , Lindsay Edwards

Charts play a vital role in data visualization, understanding data patterns, and informed decision-making. However, their unique combination of graphical elements (e.g., bars, lines) and textual components (e.g., labels, legends) poses…

计算机视觉与模式识别 · 计算机科学 2024-02-16 Fanqing Meng , Wenqi Shao , Quanfeng Lu , Peng Gao , Kaipeng Zhang , Yu Qiao , Ping Luo

Accurate surround-view depth estimation provides a competitive alternative to laser-based sensors and is essential for 3D scene understanding in autonomous driving. While empirical studies have proposed various approaches that primarily…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Weimin Liu , Wenjun Wang , Joshua H. Meng

Autoregressive models are structurally misaligned with the inherently parallel nature of geospatial understanding, forcing a rigid sequential narrative onto scenes and fundamentally hindering the generation of structured and coherent…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Jiaqi Liu , Ronghao Fu , Haoran Liu , Lang Sun , Bo Yang

Transformer-based general visual geometry frameworks have shown promising performance in camera pose estimation and 3D scene understanding. Recent advancements in Visual Geometry Grounded Transformer (VGGT) models have shown great promise…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Yangfan Xu , Lilian Zhang , Xiaofeng He , Pengdong Wu , Wenqi Wu , Jun Mao
‹ 上一页 1 2 3 10 下一页 ›