English
Related papers

Related papers: RANGER: A Monocular Zero-Shot Semantic Navigation …

200 papers

Vision-Language Navigation (VLN) is evolving from single-point pathfinding toward the more challenging Multi-Goal VLN. This task requires agents to accurately identify multiple entities while collaboratively reasoning over their…

Artificial Intelligence · Computer Science 2026-03-05 Ling Luo , Qiangian Bai

Vision-based localization in a prior map is of crucial importance for autonomous vehicles. Given a query image, the goal is to estimate the camera pose corresponding to the prior map, and the key is the registration problem of camera images…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Xingyu Chen , Jianru Xue , Shanmin Pang

In this paper, we present LOC-ZSON, a novel Language-driven Object-Centric image representation for object navigation task within complex scenes. We propose an object-centric image representation and corresponding losses for visual-language…

Computer Vision and Pattern Recognition · Computer Science 2024-05-10 Tianrui Guan , Yurou Yang , Harry Cheng , Muyuan Lin , Richard Kim , Rajasimman Madhivanan , Arnie Sen , Dinesh Manocha

Understanding natural-language references to objects in dynamic 3D driving scenes is essential for interactive autonomous systems. In practice, many referring expressions describe targets through recent motion or short-term interactions,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Jiahong Yu , Ziqi Wang , Hailiang Zhao , Wei Zhai , Xueqiang Yan , Shuiguang Deng

Modern autonomous navigation systems predominantly rely on lidar and depth cameras. However, a fundamental question remains: Can flying robots navigate in clutter using solely monocular RGB images? Given the prohibitive costs of real-world…

Robotics · Computer Science 2025-12-22 Xijie Huang , Jinhan Li , Tianyue Wu , Xin Zhou , Zhichao Han , Fei Gao

High precision localization is a crucial requirement for the autonomous driving system. Traditional positioning methods have some limitations in providing stable and accurate vehicle poses, especially in an urban environment. Herein, we…

Robotics · Computer Science 2018-05-17 Zhongyang Xiao , Kun Jiang , Shichao Xie , Tuopu Wen , Chunlei Yu , Diange Yang

Independent indoor mobility remains a critical challenge for individuals with visual impairments, largely due to the limited capability of existing assistive systems in detecting fine-grained hazardous objects such as chairs, tables, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Qi He , XiangXiang Wang , Jingtao Zhang , Yongbin Yu , Hongxiang Chu , Manping Fan , JingYe Cai , Zhenglin Yang

Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping have enabled generalizable semantic representations by…

Robotics · Computer Science 2025-08-05 Danyang Li , Zenghui Yang , Guangpeng Qi , Songtao Pang , Guangyong Shang , Qiang Ma , Zheng Yang

Localizing objects and parts from natural language in 3D space is essential for robotics, AR, and embodied AI, yet existing methods face a trade-off between the accuracy and geometric consistency of per-scene optimization and the efficiency…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Bryce Grant , Aryeh Rothenberg , Atri Banerjee , Peng Wang

3D Visual Grounding (3DVG) aims to localize objects in 3D scenes using natural language descriptions. Although supervised methods achieve higher accuracy in constrained settings, zero-shot 3DVG holds greater promise for real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Jiawen Lin , Shiran Bian , Yihang Zhu , Wenbin Tan , Yachao Zhang , Yuan Xie , Yanyun Qu

Grasping unknown objects in unstructured environments is a critical challenge for service robots, which must operate in dynamic, real-world settings such as homes, hospitals, and warehouses. Success in these environments requires both…

Robotics · Computer Science 2026-02-17 Avihai Giuili , Rotem Atari , Avishai Sintov

Extensions of Neural Radiance Fields (NeRFs) to model dynamic scenes have enabled their near photo-realistic, free-viewpoint rendering. Although these methods have shown some potential in creating immersive experiences, two drawbacks limit…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang , Pedro Miraldo , Suhas Lohit , Moitreya Chatterjee

Search and rescue operations require unmanned aerial vehicles to both traverse unknown unstructured environments at high speed and track targets once detected. Achieving both capabilities under degraded sensing and without global…

Robotics · Computer Science 2025-09-30 Alessandro Saviolo , Jeffrey Mao , Giuseppe Loianno

Autonomous scene exposure and exploration, especially in localization or communication-denied areas, useful for finding targets in unknown scenes, remains a challenging problem in computer navigation. In this work, we present a novel method…

Computer Vision and Pattern Recognition · Computer Science 2022-06-14 Tom Avrech , Evgenii Zheltonozhskii , Chaim Baskin , Ehud Rivlin

Localization is one of the most crucial tasks for Unmanned Aerial Vehicle systems (UAVs) directly impacting overall performance, which can be achieved with various sensors and applied to numerous tasks related to search and rescue…

Robotics · Computer Science 2024-11-05 Thanh Nguyen Canh , Huy-Hoang Ngo , Xiem HoangVan , Nak Young Chong

Fiducial markers can encode rich information about the environment and can aid Visual SLAM (VSLAM) approaches in reconstructing maps with practical semantic information. Current marker-based VSLAM approaches mainly utilize markers for…

Robotics · Computer Science 2023-12-27 Ali Tourani , Hriday Bavle , Jose Luis Sanchez-Lopez , Rafael Munoz Salinas , Holger Voos

Visual grounding, the task of linking textual queries to specific regions within images, plays a pivotal role in vision-language integration. Existing methods typically rely on extensive task-specific annotations and fine-tuning, limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Liqin Luo , Guangyao Chen , Xiawu Zheng , Yongxing Dai , Yixiong Zou , Yonghong Tian

Cooperative visual semantic navigation is a foundational capability for aerial robot teams operating in unknown environments. However, achieving robust open-vocabulary object-goal navigation remains challenging due to the computational…

Robotics · Computer Science 2026-03-17 MoniJesu Wonders James , Amir Atef Habel , Aleksey Fedoseev , Dzmitry Tsetserokou

The increasingly complex and diverse planetary exploration environment requires more adaptable and flexible rover navigation strategy. In this study, we propose a VLM-empowered multi-mode system to achieve efficient while safe autonomous…

Robotics · Computer Science 2025-06-23 Sinuo Cheng , Ruyi Zhou , Wenhao Feng , Huaiguang Yang , Haibo Gao , Zongquan Deng , Liang Ding

Deep Learning based techniques have been adopted with precision to solve a lot of standard computer vision problems, some of which are image classification, object detection and segmentation. Despite the widespread success of these…

Computer Vision and Pattern Recognition · Computer Science 2016-11-21 Vikram Mohanty , Shubh Agrawal , Shaswat Datta , Arna Ghosh , Vishnu Dutt Sharma , Debashish Chakravarty