中文
相关论文

相关论文: Grid-augmented vision: A simple yet effective appr…

200 篇论文

A detailed environment perception is a crucial component of automated vehicles. However, to deal with the amount of perceived information, we also require segmentation strategies. Based on a grid map environment representation, well-suited…

计算机视觉与模式识别 · 计算机科学 2018-12-06 Sascha Wirges , Tom Fischer , Jesus Balado Frias , Christoph Stiller

Alignment of large language models remains a central challenge in natural language processing. Preference optimization has emerged as a popular and effective method for improving alignment, typically through training-time or prompt-based…

机器学习 · 计算机科学 2025-10-01 Frédéric Berdoz , Luca A. Lanzendörfer , René Caky , Roger Wattenhofer

Grid cells in the entorhinal cortex encode the position of an animal in its environment using spatially periodic tuning curves of varying periodicity. Recent experiments established that these cells are functionally organized in discrete…

神经元与认知 · 定量生物学 2016-01-13 Noga Weiss Mosheiff , Haggai Agmon , Avraham Moriel , Yoram Burak

Uniform and variable environments still remain a challenge for stable visual localization and mapping in mobile robot navigation. One of the possible approaches suitable for such environments is appearance-based teach-and-repeat navigation,…

机器人学 · 计算机科学 2025-03-18 Václav Truhlařík , Tomáš Pivoňka , Michal Kasarda , Libor Přeučil

Current vision models typically maintain a fixed correspondence between their representation structure and image space. Each layer comprises a set of tokens arranged "on-the-grid," which biases patches or tokens to encode information at a…

Large foundation models trained on large-scale vision-language data can boost Open-Vocabulary Object Detection (OVD) via synthetic training data, yet the hand-crafted pipelines often introduce bias and overfit to specific prompts. We…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Yang Zhou , Shiyu Zhao , Yuxiao Chen , Zhenting Wang , Can Jin , Dimitris N. Metaxas

Grid-based structures are commonly used to encode explicit features for graphics primitives such as images, signed distance functions (SDF), and neural radiance fields (NeRF) due to their simple implementation. However, in $n$-dimensional…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Yibo Wen , Yunfan Yang

Weakly supervised visual grounding aims to predict the region in an image that corresponds to a specific linguistic query, where the mapping between the target object and query is unknown in the training stage. The state-of-the-art method…

计算机视觉与模式识别 · 计算机科学 2023-02-23 Viet-Quoc Pham , Nao Mishima

This paper introduces a novel unsupervised neural network model for visual information encoding which aims to address the problem of large-scale visual localization. Inspired by the structure of the visual cortex, the model (namely HSD)…

计算机视觉与模式识别 · 计算机科学 2021-10-01 Sylvain Colomer , Nicolas Cuperlier , Guillaume Bresson , Olivier Romain

We introduce a novel problem, i.e., the localization of an input image within a multi-modal reference map represented by a database of 3D scene graphs. These graphs comprise multiple modalities, including object-level point clouds, images,…

计算机视觉与模式识别 · 计算机科学 2024-07-15 Yang Miao , Francis Engelmann , Olga Vysotska , Federico Tombari , Marc Pollefeys , Dániel Béla Baráth

Visual navigation for autonomous agents is a core task in the fields of computer vision and robotics. Learning-based methods, such as deep reinforcement learning, have the potential to outperform the classical solutions developed for this…

计算机视觉与模式识别 · 计算机科学 2021-03-23 Zachary Seymour , Kowshik Thopalli , Niluthpol Mithun , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

Higher-dimensional spaces are ubiquitous in applications of mathematics. Yet, as we live in a three-dimensional space, visualizing, say, a four-dimensional space is challenging. We introduce a novel method of interactive visualization of…

图形学 · 计算机科学 2021-10-04 Eryk Kopczyński , Dorota Celińska-Kopczyńska

Positional encoding is important for vision transformer (ViT) to capture the spatial structure of the input image. General effectiveness has been proven in ViT. In our work we propose to train ViT to recognize the positional label of…

计算机视觉与模式识别 · 计算机科学 2023-02-17 Zhemin Zhang , Xun Gong

This study investigates the spatial reasoning capabilities of vision-language models (VLMs) through Chain-of-Thought (CoT) prompting and reinforcement learning. We begin by evaluating the impact of different prompting strategies and find…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Binbin Ji , Siddharth Agrawal , Qiance Tang , Yvonne Wu

We introduce anchored radial observations (ARO), a novel shape encoding for learning implicit field representation of 3D shapes that is category-agnostic and generalizable amid significant shape variations. The main idea behind our work is…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Yizhi Wang , Zeyu Huang , Ariel Shamir , Hui Huang , Hao Zhang , Ruizhen Hu

Vision-Language Encoders (VLEs) are widely adopted as the backbone of zero-shot referring image segmentation (RIS), enabling text-guided localization without task-specific training. However, prior works underexplored the underlying biases…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Na Min An , Inha Kang , Minhyun Lee , Hyunjung Shim

This paper presents MetricGrids, a novel grid-based neural representation that combines elementary metric grids in various metric spaces to approximate complex nonlinear signals. While grid-based representations are widely adopted for their…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Shu Wang , Yanbo Gao , Shuai Li , Chong Lv , Xun Cai , Chuankun Li , Hui Yuan , Jinglin Zhang

Object-oriented maps are important for scene understanding since they jointly capture geometry and semantics, allow individual instantiation and meaningful reasoning about objects. We introduce FroDO, a method for accurate 3D reconstruction…

计算机视觉与模式识别 · 计算机科学 2020-05-12 Kejie Li , Martin Rünz , Meng Tang , Lingni Ma , Chen Kong , Tanner Schmidt , Ian Reid , Lourdes Agapito , Julian Straub , Steven Lovegrove , Richard Newcombe

Graph neural networks (GNNs) provide a powerful and scalable solution for modeling continuous spatial data. However, they often rely on Euclidean distances to construct the input graphs. This assumption can be improbable in many real-world…

机器学习 · 计算机科学 2023-02-20 Konstantin Klemmer , Nathan Safir , Daniel B. Neill

Accurate prediction of driving scene is a challenging task due to uncertainty in sensor data, the complex behaviors of agents, and the possibility of multiple feasible futures. Existing prediction methods using occupancy grid maps primarily…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Rabbia Asghar , Lukas Rummelhard , Wenqian Liu , Anne Spalanzani , Christian Laugier