English
Related papers

Related papers: SpatialFly: Geometry-Guided Representation Alignme…

200 papers

This paper proposes VLA-AN, an efficient and onboard Vision-Language-Action (VLA) framework dedicated to autonomous drone navigation in complex environments. VLA-AN addresses four major limitations of existing large aerial navigation…

Robotics · Computer Science 2025-12-22 Yuze Wu , Mo Zhu , Xingxing Li , Yuheng Du , Yuxin Fan , Wenjun Li , Zhichao Han , Xin Zhou , Fei Gao

Today, low-altitude fixed-wing Unmanned Aerial Vehicles (UAVs) are largely limited to primitively follow user-defined waypoints. To allow fully-autonomous remote missions in complex environments, real-time environment-aware navigation is…

Vision-Language Navigation (VLN) presents a unique challenge for Large Vision-Language Models (VLMs) due to their inherent architectural mismatch: VLMs are primarily pretrained on static, disembodied vision-language tasks, which…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Jiaxing Liu , Zexi Zhang , Xiaoyan Li , Boyue Wang , Yongli Hu , Baocai Yin

Landslide monitoring and simulation play an important role in urban safety assessment and disaster prevention. Existing landslide simulation pipelines typically rely on digital elevation model and mesh-based representations, which are…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Zhenyu Liang , Jack C. P. Cheng

Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-answering (VQA) datasets may overfit dataset-specific biases, while integrating specialized…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Chun-Hsiao Yeh , Shengyi Qian , Manchen Wang , Yi Ma , Joseph Tighe , Fanyi Xiao

Embodied agents face a critical dilemma that end-to-end models lack interpretability and explicit 3D reasoning, while modular systems ignore cross-component interdependencies and synergies. To bridge this gap, we propose the Dynamic 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zihan Wang , Seungjun Lee , Guangzhao Dai , Gim Hee Lee

Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D understanding or reliance on statistical shortcuts in natural images. We introduce a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Cheolhong Min , Jaeyun Jung , Daeun Lee , Hyeonseong Jeon , Yu Su , Jonathan Tremblay , Chan Hee Song , Jaesik Park

Visual Place Recognition (vPR) plays a crucial role in Unmanned Aerial Vehicle (UAV) navigation, enabling robust localization across diverse environments. Despite significant advancements, aerial vPR faces unique challenges due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Ioannis Tsampikos Papapetros , Ioannis Kansizoglou , Antonios Gasteratos

To autonomously navigate in real-world environments, special in search and rescue operations, Unmanned Aerial Vehicles (UAVs) necessitate comprehensive maps to ensure safety. However, the prevalent metric map often lacks semantic…

Robotics · Computer Science 2024-01-17 Thanh Nguyen Canh , Armagan Elibol , Nak Young Chong , Xiem HoangVan

Geospatial sensor data is essential for modern defense and security, offering indispensable 3D information for situational awareness. This data, gathered from sources like lidar sensors and optical cameras, allows for the creation of…

Graphics · Computer Science 2025-11-10 Benjamin Kahl , Marcus Hebel , Michael Arens

We present MetaSpatial, the first reinforcement learning (RL)-based framework designed to enhance 3D spatial reasoning in vision-language models (VLMs), enabling real-time 3D scene generation without the need for hard-coded optimizations.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Zhenyu Pan , Han Liu

Vision-and-Language Navigation (VLN) has recently benefited from Multimodal Large Language Models (MLLMs), enabling zero-shot navigation. While recent exploration-based zero-shot methods have shown promising results by leveraging global…

Robotics · Computer Science 2026-03-31 Jiwen Zhang , Xiangyu Shi , Siyuan Wang , Zerui Li , Zhongyu Wei , Qi Wu

The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Thomas Monninger , Shaoyuan Xie , Qi Alfred Chen , Sihao Ding

Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current vision-language-action (VLA) models are predominantly pretrained on 2D image data without explicit 3D geometric supervision, resulting in…

We present a waypoint planning algorithm for an unmanned aerial vehicle (UAV) that is teamed with an unmanned ground vehicle (UGV) for the task of search and rescue in a subterranean environment. The UAV and UGV are teamed such that the…

Robotics · Computer Science 2021-02-12 Matteo De Petrillo , Jared Beard , Yu Gu , Jason N. Gross

Compared with existing vehicle re-identification (ReID) tasks conducted with datasets collected by fixed surveillance cameras, vehicle ReID for unmanned aerial vehicle (UAV) is still under-explored and could be more challenging. Vehicles…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Aihuan Yao , Jiahao Qi , Ping Zhong

Absolute Visual Localization (AVL) enables an Unmanned Aerial Vehicle (UAV) to determine its position in GNSS-denied environments by establishing geometric relationships between UAV images and geo-tagged reference maps. While many previous…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Yibin Ye , Xichao Teng , Shuo Chen , Leqi Liu , Kun Wang , Xiaokai Song , Zhang Li

While current multimodal models can answer questions based on 2D images, they lack intrinsic 3D object perception, limiting their ability to comprehend spatial relationships and depth cues in 3D scenes. In this work, we propose N3D-VLM, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Yuxin Wang , Lei Ke , Boqiang Zhang , Tianyuan Qu , Hanxun Yu , Zhenpeng Huang , Meng Yu , Dan Xu , Dong Yu

Precise geolocalization is crucial for unmanned aerial vehicles (UAVs). However, most current deployed UAVs rely on the global navigation satellite systems (GNSS) or high precision inertial navigation systems (INS) for geolocalization. In…

Robotics · Computer Science 2023-01-02 Jun Mao , Lilian Zhang , Xiaofeng He , Hao Qu , Xiaoping Hu

Current state-of-the-art spatial reasoning-enhanced VLMs are trained to excel at spatial visual question answering (VQA). However, we believe that higher-level 3D-aware tasks, such as articulating dynamic scene changes and motion planning,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Chenyang Ma , Kai Lu , Ta-Ying Cheng , Niki Trigoni , Andrew Markham