中文
相关论文

相关论文: G^3: Geolocation via Guidebook Grounding

200 篇论文

Visual Geo-localization (VG) refers to the process to identify the location described in query images, which is widely applied in robotics field and computer vision tasks, such as autonomous driving, metaverse, augmented reality, and SLAM.…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Chen Mao , Jingqi Hu

We propose a simple yet effective text- based user geolocation model based on a neural network with one hidden layer, which achieves state of the art performance over three Twitter benchmark geolocation datasets, in addition to producing…

计算与语言 · 计算机科学 2017-04-28 Afshin Rahimi , Trevor Cohn , Timothy Baldwin

Deep learning has shown strong performance in geospatial prediction tasks, but the role of geolocation information in improving accuracy and generalizability remains underexamined. Recent work has introduced location encoders that aim to…

机器学习 · 计算机科学 2025-10-28 Morteza Karimzadeh , Zhongying Wang , James L. Crooks

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing…

机器人学 · 计算机科学 2026-02-05 Rui Tang , Guankun Wang , Long Bai , Huxin Gao , Jiewen Lai , Chi Kit Ng , Jiazheng Wang , Fan Zhang , Hongliang Ren

Transformer-based general visual geometry frameworks have shown promising performance in camera pose estimation and 3D scene understanding. Recent advancements in Visual Geometry Grounded Transformer (VGGT) models have shown great promise…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Yangfan Xu , Lilian Zhang , Xiaofeng He , Pengdong Wu , Wenqi Wu , Jun Mao

Text data are an important source of detailed information about social and political events. Automated systems parse large volumes of text data to infer or extract structured information that describes actors, actions, dates, times, and…

计算与语言 · 计算机科学 2021-07-02 Benjamin J. Radford

This paper presents a localization algorithm for autonomous urban vehicles under rain weather conditions. In adverse weather, human drivers anticipate the location of the ego-vehicle based on the control inputs they provide and surrounding…

机器人学 · 计算机科学 2023-06-16 Yu Xiang Tan , Malika Meghjani , Marcel Bartholomeus Prasetyo

Many applications such as autonomous navigation, urban planning and asset monitoring, rely on the availability of accurate information about objects and their geolocations. In this paper we propose to automatically detect and compute the…

计算机视觉与模式识别 · 计算机科学 2018-05-08 Vladimir A. Krylov , Eamonn Kenny , Rozenn Dahyot

Machine translation between many languages at once is highly challenging, since training with ground truth requires supervision between all language pairs, which is difficult to obtain. Our key insight is that, while languages may vary…

计算与语言 · 计算机科学 2022-04-04 Dídac Surís , Dave Epstein , Carl Vondrick

Understanding how AI will represent and reason about geography should be a key concern for all of us, as the broader public increasingly interacts with spaces and places through these systems. Similarly, in line with the nature of…

人工智能 · 计算机科学 2026-03-20 Krzysztof Janowicz , Gengchen Mai , Rui Zhu , Song Gao , Zhangyu Wang , Yingjie Hu , Lauren Bennett

We tackle the problem of finding accurate and robust keypoint correspondences between images. We propose a learning-based approach to guide local feature matches via a learned approximate image matching. Our approach can boost the results…

计算机视觉与模式识别 · 计算机科学 2021-05-03 François Darmon , Mathieu Aubry , Pascal Monasse

The goal of cross-view image based geo-localization is to determine the location of a given street view image by matching it against a collection of geo-tagged satellite images. This task is notoriously challenging due to the drastic…

计算机视觉与模式识别 · 计算机科学 2021-03-12 Aysim Toker , Qunjie Zhou , Maxim Maximov , Laura Leal-Taixé

Images are a convenient way to specify which particular object instance an embodied agent should navigate to. Solving this task requires semantic visual reasoning and exploration of unknown environments. We present a system that can perform…

In this paper, we study Tracking by Language that localizes the target box sequence in a video based on a language query. We propose a framework called GTI that decomposes the problem into three sub-tasks: Grounding, Tracking, and…

计算机视觉与模式识别 · 计算机科学 2020-11-17 Zhengyuan Yang , Tushar Kumar , Tianlang Chen , Jinsong Su , Jiebo Luo

Contrastive learning methods have significantly narrowed the gap between supervised and unsupervised learning on computer vision tasks. In this paper, we explore their application to geo-located datasets, e.g. remote sensing, where…

计算机视觉与模式识别 · 计算机科学 2022-03-09 Kumar Ayush , Burak Uzkent , Chenlin Meng , Kumar Tanmay , Marshall Burke , David Lobell , Stefano Ermon

Visual-Language Models (VLMs) have shown remarkable performance across various tasks, particularly in recognizing geographic information from images. However, VLMs still show regional biases in this task. To systematically evaluate these…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Jingyuan Huang , Jen-tse Huang , Ziyi Liu , Xiaoyuan Liu , Wenxuan Wang , Jieyu Zhao

Localizing textual descriptions within large-scale 3D scenes presents inherent ambiguities, such as identifying all traffic lights in a city. Addressing this, we introduce a method to generate distributions of camera poses conditioned on…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Qi Ma , Runyi Yang , Bin Ren , Nicu Sebe , Ender Konukoglu , Luc Van Gool , Danda Pani Paudel

The task of cross-view image geo-localization aims to determine the geo-location (GPS coordinates) of a query ground-view image by matching it with the GPS-tagged aerial (satellite) images in a reference dataset. Due to the dramatic changes…

计算机视觉与模式识别 · 计算机科学 2019-04-15 Bin Sun , Chen Chen , Yingying Zhu , Jianmin Jiang

Pretrained language models (PLMs) often fail to fairly represent target users from certain world regions because of the under-representation of those regions in training datasets. With recent PLMs trained on enormous data sources,…

计算与语言 · 计算机科学 2022-12-21 Fahim Faisal , Antonios Anastasopoulos

Robots that can manipulate objects in unstructured environments and collaborate with humans can benefit immensely by understanding natural language. We propose a pipelined architecture of two stages to perform spatial reasoning on the text…