English
Related papers

Related papers: Grid-augmented vision: A simple yet effective appr…

200 papers

We propose a novel method for geolocalizing Unmanned Aerial Vehicles (UAVs) in environments lacking Global Navigation Satellite Systems (GNSS). Current state-of-the-art techniques employ an offline-trained encoder to generate a vector…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Theo Di Piazza , Enric Meinhardt-Llopis , Gabriele Facciolo , Benedicte Bascle , Corentin Abgrall , Jean-Clement Devaux

We propose VISO-Grasp, a novel vision-language-informed system designed to systematically address visibility constraints for grasping in severely occluded environments. By leveraging Foundation Models (FMs) for spatial reasoning and active…

Robotics · Computer Science 2025-08-07 Yitian Shi , Di Wen , Guanqi Chen , Edgar Welte , Sheng Liu , Kunyu Peng , Rainer Stiefelhagen , Rania Rayyes

GUI grounding, the task of mapping natural-language instructions to pixel coordinates, is crucial for autonomous agents, yet remains difficult for current VLMs. The core bottleneck is reliable patch-to-pixel mapping, which breaks when…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Suyuchen Wang , Tianyu Zhang , Ahmed Masry , Christopher Pal , Spandana Gella , Bang Liu , Perouz Taslakian

We propose an object detector for top-view grid maps which is additionally trained to generate an enriched version of its input. Our goal in the joint model is to improve generalization by regularizing towards structural knowledge in form…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Sascha Wirges , Ye Yang , Sven Richter , Haohao Hu , Christoph Stiller

Augmented reality (AR) offers promising opportunities to support movement-based activities, such as personal training or physical therapy, with real-time, spatially-situated visual cues. While many approaches leverage AR to guide motion,…

Human-Computer Interaction · Computer Science 2025-10-02 Jade Kandel , Sriya Kasumarthi , Spiros Tsalikis , Chelsea Duppen , Daniel Szafir , Michael Lewek , Henry Fuchs , Danielle Szafir

Visual-to-auditory sensory substitution devices can assist the blind in sensing the visual environment by translating the visual information into a sound pattern. To improve the translation quality, the task performances of the blind are…

Computer Vision and Pattern Recognition · Computer Science 2019-04-22 Di Hu , Dong Wang , Xuelong Li , Feiping Nie , Qi Wang

Evidential grids have recently shown interesting properties for mobile object perception. Evidential grids are a generalisation of Bayesian occupancy grids using Dempster- Shafer theory. In particular, these grids can handle efficiently…

Robotics · Computer Science 2014-01-23 Marek Kurdej , Julien Moras , Véronique Cherfaoui , Philippe Bonnifait

A method of information transmission using visual markers has been widely studied. In this approach, information or identifiers (IDs) are encoded in the black-and-white pattern of each marker. By analyzing the geometric properties of the…

Information Theory · Computer Science 2026-01-13 Wataru Uemura , Shogo Kawasaki

Autonomous agents such as cars, robots and drones need to precisely localize themselves in diverse environments, including in GPS-denied indoor environments. One approach for precise localization is visual place recognition (VPR), which…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Ni Wang , Zihan You , Emre Neftci , Thorben Schoepe

We propose Guided Zoom, an approach that utilizes spatial grounding of a model's decision to make more informed predictions. It does so by making sure the model has "the right reasons" for a prediction, defined as reasons that are coherent…

Computer Vision and Pattern Recognition · Computer Science 2020-03-24 Sarah Adel Bargal , Andrea Zunino , Vitali Petsiuk , Jianming Zhang , Kate Saenko , Vittorio Murino , Stan Sclaroff

Grid cells enable the brain to model the physical space of the world and navigate effectively via path integration, updating self-position using information from self-movement. Recent proposals suggest that the brain might use similar…

Artificial Intelligence · Computer Science 2021-02-19 Niels Leadholm , Marcus Lewis , Subutai Ahmad

In modern computer vision, images are typically represented as a fixed uniform grid with some stride and processed via a deep convolutional neural network. We argue that deforming the grid to better align with the high-frequency image…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Jun Gao , Zian Wang , Jinchen Xuan , Sanja Fidler

Face inpainting requires the model to have a precise global understanding of the facial position structure. Benefiting from the powerful capabilities of deep learning backbones, recent works in face inpainting have achieved decent…

Computer Vision and Pattern Recognition · Computer Science 2024-01-22 Bo Zhao , Huan Yang , Jianlong Fu

A common need for artificial intelligence models in the broader geoscience is to represent and encode various types of spatial data, such as points (e.g., points of interest), polylines (e.g., trajectories), polygons (e.g., administrative…

Machine Learning · Computer Science 2022-03-14 Gengchen Mai , Krzysztof Janowicz , Yingjie Hu , Song Gao , Bo Yan , Rui Zhu , Ling Cai , Ni Lao

High-dimensional neural activity often reside in a low-dimensional subspace, referred to as neural manifolds. Grid cells in the medial entorhinal cortex provide a periodic spatial code that are organized near a toroidal manifold,…

Neurons and Cognition · Quantitative Biology 2025-10-22 Yuxing Jared Yao , Iris H. R. Yoon

Transformer models are permutation equivariant. To supply the order and type information of the input tokens, position and segment embeddings are usually added to the input. Recent works proposed variations of positional encodings with…

Computation and Language · Computer Science 2021-11-04 Pu-Chin Chen , Henry Tsai , Srinadh Bhojanapalli , Hyung Won Chung , Yin-Wen Chang , Chun-Sung Ferng

Intelligent assistive systems can navigate blind people, but most of them could only give non-intuitive cues or inefficient guidance. Based on computer vision and vibrotactile encoding, this paper presents an interactive system that…

Robotics · Computer Science 2022-06-22 Zhikai Wei , Xuhui Hu

Existing visual localization methods are typically either 2D image-based, which are easy to build and maintain but limited in effective geometric reasoning, or 3D structure-based, which achieve high accuracy but require a centralized…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Xudong Jiang , Fangjinhua Wang , Silvano Galliani , Christoph Vogel , Marc Pollefeys

Decoding visual images from brain activity has significant potential for advancing brain-computer interaction and enhancing the understanding of human perception. Recent approaches align the representation spaces of images and brain…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Nona Rajabi , Antônio H. Ribeiro , Miguel Vasco , Farzaneh Taleb , Mårten Björkman , Danica Kragic

Visual perception is critically influenced by the focus of attention. Due to limited resources, it is well known that neural representations are biased in favor of attended locations. Using concurrent eye-tracking and functional Magnetic…

Computer Vision and Pattern Recognition · Computer Science 2020-10-02 Meenakshi Khosla , Gia H. Ngo , Keith Jamison , Amy Kuceyeski , Mert R. Sabuncu