English
Related papers

Related papers: Poly2Vec: Polymorphic Fourier-Based Encoding of Ge…

200 papers

Visual grounding is a task to locate the target indicated by a natural language expression. Existing methods extend the generic object detection framework to this problem. They base the visual grounding on the features from pre-generated…

Computer Vision and Pattern Recognition · Computer Science 2022-06-09 Li Yang , Yan Xu , Chunfeng Yuan , Wei Liu , Bing Li , Weiming Hu

Worldwide visual geo-localization aims to determine the geographic location of an image anywhere on Earth using only its visual content. Despite recent progress, learning expressive representations of geographic space remains challenging…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Angel Daruna , Nicholas Meegan , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

The choice of data representation is a key factor in the success of deep learning in geometric tasks. For instance, DUSt3R recently introduced the concept of viewpoint-invariant point maps, generalizing depth prediction and showing that all…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Ben Kaye , Tomas Jakab , Shangzhe Wu , Christian Rupprecht , Andrea Vedaldi

Many real-world physics and engineering problems arise in geometrically complex domains discretized by meshes for numerical simulations. The nodes of these potentially irregular meshes naturally form point clouds whose limited tractability…

Machine Learning · Computer Science 2025-06-17 Shirin Hosseinmardi , Ramin Bostanabad

Understanding how neural representations respond to geometric transformations is essential for evaluating whether learned features preserve meaningful spatial structure. Existing approaches primarily assess robustness primarily by comparing…

Machine Learning · Computer Science 2026-05-12 Huahua Lin , Katayoun Farrahi , Xiaohao Cai

The Earth's surface is continually changing, and identifying changes plays an important role in urban planning and sustainability. Although change detection techniques have been successfully developed for many years, these techniques are…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Zhenghang Yuan , Lichao Mou , Zhitong Xiong , Xiaoxiang Zhu

We introduce a novel task of 3D visual grounding in monocular RGB images using language descriptions with both appearance and geometry information. Specifically, we build a large-scale dataset, Mono3DRefer, which contains 3D object targets…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Yang Zhan , Yuan Yuan , Zhitong Xiong

Multimodal large language models (MLLMs) have made significant advancements in vision understanding and reasoning. However, the autoregressive Transformer architecture used by MLLMs requries tokenization on input images, which limits their…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Xiangxuan Ren , Zhongdao Wang , Liping Hou , Pin Tang , Guoqing Wang , Chao Ma

We propose DoE2Vec, a variational autoencoder (VAE)-based methodology to learn optimization landscape characteristics for downstream meta-learning tasks, e.g., automated selection of optimization algorithms. Principally, using large…

Optimization and Control · Mathematics 2023-04-05 Bas van Stein , Fu Xing Long , Moritz Frenzel , Peter Krause , Markus Gitterle , Thomas Bäck

Multimodal large language models (MLLMs) have achieved significant progress in image and language tasks due to the strong reasoning capability of large language models (LLMs). Nevertheless, most MLLMs suffer from limited spatial reasoning…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Jiajie Guo , Qingpeng Zhu , Jin Zeng , Xiaolong Wu , Changyong He , Weida Wang

Vision-language models (VLMs) have advanced multimodal reasoning but still face challenges in spatial reasoning for 3D scenes and complex object configurations. To address this, we introduce SpatialViLT, an enhanced VLM that integrates…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Chashi Mahiul Islam , Oteo Mamo , Samuel Jacob Chacko , Xiuwen Liu , Weikuan Yu

Elevation maps are commonly used to represent the environment of mobile robots and are instrumental for locomotion and navigation tasks. However, pure geometric information is insufficient for many field applications that require appearance…

Robotics · Computer Science 2024-10-28 Gian Erni , Jonas Frey , Takahiro Miki , Matias Mattamala , Marco Hutter

Efficient point cloud coding has become increasingly critical for multiple applications such as virtual reality, autonomous driving, and digital twin systems, where rich and interactive 3D data representations may functionally make the…

Image and Video Processing · Electrical Eng. & Systems 2025-03-13 André F. R. Guarda , Nuno M. M. Rodrigues , Fernando Pereira

Reconstructing geometry and topology structures from raw unstructured data has always been an important research topic in indoor mapping research. In this paper, we aim to reconstruct the floorplan with a vectorized representation from…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Yuzhou Liu , Lingjie Zhu , Xiaodong Ma , Hanqiao Ye , Xiang Gao , Xianwei Zheng , Shuhan Shen

Object encoding and identification are crucial for many robotic tasks such as autonomous exploration and semantic relocalization. Existing works heavily rely on the tracking of detected objects but have difficulty recalling revisited…

Robotics · Computer Science 2022-01-27 Kuan Xu , Chen Wang , Chao Chen , Wei Wu , Sebastian Scherer

As a core task in location-based services (LBS) (e.g., navigation maps), query and point of interest (POI) matching connects users' intent with real-world geographic information. Recently, pre-trained models (PTMs) have made advancements in…

Computation and Language · Computer Science 2023-05-25 Ruixue Ding , Boli Chen , Pengjun Xie , Fei Huang , Xin Li , Qiang Zhang , Yao Xu

Weakly supervised localization aims at finding target object regions using only image-level supervision. However, localization maps extracted from classification networks are often not accurate due to the lack of fine pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2020-08-13 Xiaolin Zhang , Yunchao Wei , Yi Yang

In this paper, we address the problem of 6-DoF object pose estimation from a single RGB image. Indirect methods that typically predict intermediate 2D keypoints, followed by a Perspective-n-Point solver, have shown great performance. Direct…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Nassim Ali Ousalah , Peyman Rostami , Vincent Gaudillière , Emmanuel Koumandakis , Anis Kacem , Enjie Ghorbel , Djamila Aouada

3D occupancy perception holds a pivotal role in recent vision-centric autonomous driving systems by converting surround-view images into integrated geometric and semantic representations within dense 3D grids. Nevertheless, current models…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Xin Tan , Wenbin Wu , Zhiwei Zhang , Chaojie Fan , Yong Peng , Zhizhong Zhang , Yuan Xie , Lizhuang Ma

Mathematical reasoning remains an ongoing challenge for AI models, especially for geometry problems that require both linguistic and visual signals. As the vision encoders of most MLLMs are trained on natural scenes, they often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Zeren Zhang , Jo-Ku Cheng , Jingyang Deng , Lu Tian , Jinwen Ma , Ziran Qin , Xiaokai Zhang , Na Zhu , Tuo Leng
‹ Prev 1 8 9 10 Next ›