English
Related papers

Related papers: SemNav: A Model-Based Planner for Zero-Shot Object…

200 papers

Autonomous navigation based on precise localization has been widely developed in both academic research and practical applications. The high demand for localization accuracy has been essential for safe robot planing and navigation while it…

Robotics · Computer Science 2019-06-07 Huifang Ma , Yue Wang , Li Tang , Sarath Kodagoda , Rong Xiong

How can we build general-purpose robot systems for open-world semantic navigation, e.g., searching a novel environment for a target object specified in natural language? To tackle this challenge, we introduce OSG Navigator, a modular system…

Robotics · Computer Science 2025-08-07 Joel Loo , Zhanxin Wu , David Hsu

The generalization of the end-to-end deep reinforcement learning (DRL) for object-goal visual navigation is a long-standing challenge since object classes and placements vary in new test environments. Learning domain-independent visual…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Shiwei Lian , Feitian Zhang

Zero-shot Vision-and-Language Navigation (VLN) agents leveraging Large Language Models (LLMs) excel in generalization but suffer from insufficient spatial perception. Focusing on complex continuous environments, we categorize key perceptual…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Lu Yue , Yue Fan , Shiwei Lian , Yu Zhao , Jiaxin Yu , Liang Xie , Feitian Zhang

Zero-shot classification is a promising paradigm to solve an applicable problem when the training classes and test classes are disjoint. Achieving this usually needs experts to externalize their domain knowledge by manually specifying a…

Human-Computer Interaction · Computer Science 2021-08-17 Shichao Jia , Zeyu Li , Nuo Chen , Jiawan Zhang

Reinforcement learning and planning methods require an objective or reward function that encodes the desired behavior. Yet, in practice, there is a wide range of scenarios where an objective is difficult to provide programmatically, such as…

Machine Learning · Computer Science 2018-10-02 Annie Xie , Avi Singh , Sergey Levine , Chelsea Finn

Path planning is a critical component in autonomous drone operations, enabling safe and efficient navigation through complex environments. Recent advances in foundation models, particularly large language models (LLMs) and vision-language…

Robotics · Computer Science 2025-05-28 Jiaping Xiao , Cheng Wen Tsao , Yuhang Zhang , Mir Feroskhan

In visual semantic navigation, the robot navigates to a target object with egocentric visual observations and the class label of the target is given. It is a meaningful task inspiring a surge of relevant research. However, most of the…

Artificial Intelligence · Computer Science 2021-09-21 Xinzhu Liu , Di Guo , Huaping Liu , Fuchun Sun

Zero-shot scene understanding in real-world settings presents major challenges due to the complexity and variability of natural scenes, where models must recognize new objects, actions, and contexts without prior labeled examples. This work…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Manjunath Prasad Holenarasipura Rajiv , B. M. Vidyavathi

Visual navigation takes inspiration from humans, who navigate in previously unseen environments using vision without detailed environment maps. Inspired by this, we introduce a novel no-RL, no-graph, no-odometry approach to visual…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Faith Johnson , Bryan Bo Cao , Ashwin Ashok , Shubham Jain , Kristin Dana

This paper introduces a novel semantics-aware inspection planning policy derived through deep reinforcement learning. Reflecting the fact that within autonomous informative path planning missions in unknown environments, it is often only a…

Robotics · Computer Science 2025-05-21 Grzegorz Malczyk , Mihir Kulkarni , Kostas Alexis

Zero-shot object navigation in unknown environments presents significant challenges, mainly due to two key limitations: insufficient semantic guidance leads to inefficient exploration, while limited spatial memory resulting from…

Robotics · Computer Science 2025-09-30 Xiangyi Meng , Delun Li , Zihao Mao , Yi Yang , Wenjie Song

Machine learning has been considered a promising approach for indoor localization. Nevertheless, the sample efficiency, scalability, and generalization ability remain open issues of implementing learning-based algorithms in practical…

Signal Processing · Electrical Eng. & Systems 2024-05-24 Haiyao Yu , Changyang She , Yunkai Hu , Geng Wang , Rui Wang , Branka Vucetic , Yonghui Li

Breakthrough progress in vision-based navigation through unknown environments has been achieved by using multimodal large language models (MLLMs). These models can plan a sequence of motions by evaluating the current view at each time step…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Wanrong Zheng , Yunhao Ge , Laurent Itti

This work studies object goal navigation task, which involves navigating to the closest object related to the given semantic category in unseen environments. Recent works have shown significant achievements both in the end-to-end…

Artificial Intelligence · Computer Science 2021-09-21 Aleksey Staroverov , Aleksandr I. Panov

Embodied agents equipped with GPT as their brains have exhibited extraordinary decision-making and generalization abilities across various tasks. However, existing zero-shot agents for vision-and-language navigation (VLN) only prompt GPT-4…

Artificial Intelligence · Computer Science 2024-06-21 Jiaqi Chen , Bingqian Lin , Ran Xu , Zhenhua Chai , Xiaodan Liang , Kwan-Yee K. Wong

The ability to accurately locate and navigate to a specific object is a crucial capability for embodied agents that operate in the real world and interact with objects to complete tasks. Such object navigation tasks usually require…

Artificial Intelligence · Computer Science 2023-07-07 Kaiwen Zhou , Kaizhi Zheng , Connor Pryor , Yilin Shen , Hongxia Jin , Lise Getoor , Xin Eric Wang

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, which requires multi-round dialogue spatial reasoning and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Xunyi Zhao , Gengze Zhou , Qi Wu

The advances in deep reinforcement learning recently revived interest in data-driven learning based approaches to navigation. In this paper we propose to learn viewpoint invariant and target invariant visual servoing for local mobile robot…

Computer Vision and Pattern Recognition · Computer Science 2020-03-06 Yimeng Li , Jana Kosecka

Semantic reasoning and dynamic planning capabilities are crucial for an autonomous agent to perform complex navigation tasks in unknown environments. It requires a large amount of common-sense knowledge, that humans possess, to succeed in…

Robotics · Computer Science 2024-04-05 Abhinav Rajvanshi , Karan Sikka , Xiao Lin , Bhoram Lee , Han-Pang Chiu , Alvaro Velasquez