English
Related papers

Related papers: Where Do We Go from Here? Multi-scale Allocentric …

200 papers

Target-driven visual navigation aims at navigating an agent towards a given target based on the observation of the agent. In this task, it is critical to learn informative visual representation and robust navigation policy. Aiming to…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Heming Du , Xin Yu , Liang Zheng

Variation in spatial categorization across languages is often studied by eliciting human labels for the relations depicted in a set of scenes known as the Topological Relations Picture Series (TRPS). We demonstrate that labels generated by…

Computation and Language · Computer Science 2026-03-11 Wanchun Li , Alexandra Carstensen , Yang Xu , Terry Regier , Charles Kemp

Real-world applications, such as autonomous driving and humanoid robot manipulation, require precise spatial perception. However, it remains underexplored how Vision-Language Models (VLMs) recognize spatial relationships and perceive…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Fei Kong , Jinhao Duan , Kaidi Xu , Zhenhua Guo , Xiaofeng Zhu , Xiaoshuang Shi

Many place-related questions can only be answered by complex spatial reasoning, a task poorly supported by factoid question retrieval. Such reasoning using combinations of spatial and non-spatial criteria pertinent to place-related…

Information Retrieval · Computer Science 2022-05-09 Ehsan Hamzei , Martin Tomko , Stephan Winter

We introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset. RxR is multilingual (English, Hindi, and Telugu) and larger (more paths and instructions) than other VLN datasets. It emphasizes the role of language…

Computer Vision and Pattern Recognition · Computer Science 2020-10-19 Alexander Ku , Peter Anderson , Roma Patel , Eugene Ie , Jason Baldridge

We improve reliable, long-horizon, goal-directed navigation in partially-mapped environments by using non-locally available information to predict the goodness of temporally-extended actions that enter unseen space. Making predictions about…

Robotics · Computer Science 2024-03-08 Raihan Islam Arnob , Gregory J. Stein

Geo-localization aims to infer the geographic location where an image was captured using observable visual evidence. Traditional methods achieve impressive results through large-scale training on massive image corpora. With the emergence of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Jinnao Li , Zijian Chen , Tingzhu Chen , Changbo Wang

Human beings cooperatively navigate rule-constrained environments by adhering to mutually known navigational patterns, which may be represented as directional pathways or road lanes. Inferring these navigational patterns from incompletely…

Computer Vision and Pattern Recognition · Computer Science 2023-07-07 Robin Karlsson , Alexander Carballo , Francisco Lepe-Salazar , Keisuke Fujii , Kento Ohtani , Kazuya Takeda

The hippocampal formation is thought to learn spatial maps of environments, and in many models this learning process consists of forming a sensory association for each location in the environment. This is inefficient, akin to learning a…

Artificial Intelligence · Computer Science 2021-07-02 Marcus Lewis

Navigation aids are central to immersive virtual reality (VR) experiences that involve physical locomotion. Their effectiveness depends not only on how much spatial information they provide, but also on how directly that information…

Human-Computer Interaction · Computer Science 2026-03-19 Apurv Varshney , Lily M. Turkstra , Jiaxin Su , Mable Zhou , Scott T. Grafton , Barry Giesbrecht , Mary Hegarty , Michael Beyeler

We introduce a dataset containing human-authored descriptions of target locations in an "end-of-trip in a taxi ride" scenario. We describe our data collection method and a novel annotation scheme that supports understanding of such…

Computation and Language · Computer Science 2018-07-12 Deepthi Karkada , Ramesh Manuvinakurike , Kallirroi Georgila

Humans have a rich representation of the entities in their environment. Entities are described by their attributes, and entities that share attributes are often semantically related. For example, if two books have "Natural Language…

Artificial Intelligence · Computer Science 2020-11-23 Mohamadreza Faridghasemnia , Daniele Nardi , Alessandro Saffiotti

This paper explores leveraging large language models for map-free off-road navigation using generative AI, reducing the need for traditional data collection and annotation. We propose a method where a robot receives verbal instructions,…

Robotics · Computer Science 2024-04-04 Faraz Lotfi , Farnoosh Faraji , Nikhil Kakodkar , Travis Manderson , David Meger , Gregory Dudek

Human navigation has been a topic of interest in spatial cognition from the past few decades. It has been experimentally observed that humans accomplish the task of way-finding a destination in an unknown environment by recognizing…

Social and Information Networks · Computer Science 2011-11-22 Vijesh M. , Sudarshan Iyengar , Vijay Mahantesh , Amitash Ramesh , Veni Madhavan

Humans are known to construct cognitive maps of their everyday surroundings using a variety of perceptual inputs. As such, when a human is asked for directions to a particular location, their wayfinding capability in converting this…

Robotics · Computer Science 2020-11-03 Vishnu Sashank Dorbala , Arjun Srinivasan , Aniket Bera

Geospatial reasoning is essential for real-world applications such as urban analytics, transportation planning, and disaster response. However, existing LLM-based agents often fail at genuine geospatial computation, relying instead on web…

Artificial Intelligence · Computer Science 2026-01-26 Riyang Bao , Cheng Yang , Dazhou Yu , Zhexiang Tang , Gengchen Mai , Liang Zhao

Training a reinforcement learning agent to carry out natural language instructions is limited by the available supervision, i.e. knowing when the instruction has been carried out. We adapt the CLEVR visual question answering dataset to…

Machine Learning · Computer Science 2021-06-04 Michiel de Jong , Satyapriya Krishna , Anuva Agarwal

We study the task of semantic mapping - specifically, an embodied agent (a robot or an egocentric AI assistant) is given a tour of a new environment and asked to build an allocentric top-down semantic map ("what is where?") from egocentric…

Computer Vision and Pattern Recognition · Computer Science 2021-03-12 Vincent Cartillier , Zhile Ren , Neha Jain , Stefan Lee , Irfan Essa , Dhruv Batra

State-of-the-art navigation methods leverage a spatial memory to generalize to new environments, but their occupancy maps are limited to capturing the geometric structures directly observed by the agent. We propose occupancy anticipation,…

Computer Vision and Pattern Recognition · Computer Science 2020-08-26 Santhosh K. Ramakrishnan , Ziad Al-Halah , Kristen Grauman

Visual Place Recognition is a task that aims to predict the place of an image (called query) based solely on its visual features. This is typically done through image retrieval, where the query is matched to the most similar images from a…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Gabriele Berton , Gabriele Trivigno , Barbara Caputo , Carlo Masone