English
Related papers

Related papers: SFCo-Nav: Efficient Zero-Shot Visual Language Navi…

200 papers

With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream tasks. However, the requirement for expensive annotations…

Computer Vision and Pattern Recognition · Computer Science 2022-05-17 Zirui Wang , Jiahui Yu , Adams Wei Yu , Zihang Dai , Yulia Tsvetkov , Yuan Cao

The Zero-Shot Object Navigation (ZSON) task requires embodied agents to find a previously unseen object by navigating in unfamiliar environments. Such a goal-oriented exploration heavily relies on the ability to perceive, understand, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Linqing Zhong , Chen Gao , Zihan Ding , Yue Liao , Huimin Ma , Shifeng Zhang , Xu Zhou , Si Liu

We study the task of zero-shot vision-and-language navigation (ZS-VLN), a practical yet challenging problem in which an agent learns to navigate following a path described by language instructions without requiring any path-instruction…

Computer Vision and Pattern Recognition · Computer Science 2023-08-17 Peihao Chen , Xinyu Sun , Hongyan Zhi , Runhao Zeng , Thomas H. Li , Gaowen Liu , Mingkui Tan , Chuang Gan

Indoor mobile robot navigation requires fast responsiveness and robust semantic understanding, yet existing methods struggle to provide both. Classical geometric approaches such as SLAM offer reliable localization but depend on detailed…

Robotics · Computer Science 2026-01-30 Joonhee Lee , Hyunseung Shin , Jeonggil Ko

For robots navigating in human-populated environments, safety and social compliance are equally critical, yet prior work has mostly emphasized safety. Socially compliant navigation that accounts for human comfort, social norms, and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Tomohito Kawabata , Xinyu Zhang , Ling Xiao

Despite the impressive advancements of Large Vision-Language Models (LVLMs), existing approaches suffer from a fundamental bottleneck: inefficient visual-language integration. Current methods either disrupt the model's inherent structure or…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Tongtian Yue , Longteng Guo , Yepeng Tang , Zijia Zhao , Xinxin Zhu , Hua Huang , Jing Liu

Vision-and-Language Navigation (VLN) in continuous environments requires agents to interpret natural language instructions while navigating unconstrained 3D spaces. Existing VLN-CE frameworks rely on a two-stage approach: a waypoint…

Robotics · Computer Science 2025-06-18 Xiangyu Shi , Zerui Li , Wenqi Lyu , Jiatong Xia , Feras Dayoub , Yanyuan Qiao , Qi Wu

Existing aerial Vision-Language Navigation (VLN) methods predominantly adopt a detection-and-planning pipeline, which converts open-vocabulary detections into discrete textual scene graphs. These approaches are plagued by inadequate spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Haoyu Tong , Xiangyu Dong , Xiaoguang Ma , Haoran Zhao , Yaoming Zhou , Chenghao Lin

Vision-language models (VLMs) pre-trained on natural image and language data, such as CLIP, have exhibited significant potential in few-shot image recognition tasks, leading to development of various efficient transfer learning methods.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Dexia Chen , Wentao Zhang , Qianjie Zhu , Ping Hu , Weibing Li , Tong Zhang , Ruixuan Wang

Vision-Language Models (VLMs) achieve strong performance on spatial question answering benchmarks, yet it remains unclear whether such gains reflect genuine spatial intelligence. We show that existing spatial VLMs lack basic camera motion…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Hsiang-Wei Huang , Junbin Lu , Kuang-Ming Chen , Jianxu Shangguan , Cheng-Yen Yang , Jenq-Neng Hwang

We propose VLM-Social-Nav, a novel Vision-Language Model (VLM) based navigation approach to compute a robot's motion in human-centered environments. Our goal is to make real-time decisions on robot actions that are socially compliant with…

Robotics · Computer Science 2024-11-27 Daeun Song , Jing Liang , Amirreza Payandeh , Amir Hossain Raj , Xuesu Xiao , Dinesh Manocha

Maritime port inspection plays a critical role in ensuring safety, regulatory compliance, and operational efficiency in complex maritime environments. However, existing inspection methods often rely on manual operations and conventional…

Robotics · Computer Science 2026-01-21 Muhayy Ud Din , Waseem Akram , Ahsan B. Bakht , Irfan Hussain

Object navigation is a core capability of embodied intelligence, enabling an agent to locate target objects in unknown environments. Recent advances in vision-language models (VLMs) have facilitated zero-shot object navigation (ZSON).…

Robotics · Computer Science 2026-02-13 Wancai Zheng , Hao Chen , Xianlong Lu , Linlin Ou , Xinyi Yu

We present VLMnav, an embodied framework to transform a Vision-Language Model (VLM) into an end-to-end navigation policy. In contrast to prior work, we do not rely on a separation between perception, planning, and control; instead, we use a…

Robotics · Computer Science 2024-11-11 Dylan Goetting , Himanshu Gaurav Singh , Antonio Loquercio

Vision-language navigation (VLN) has emerged as a promising paradigm, enabling mobile robots to perform zero-shot inference and execute tasks without specific pre-programming. However, current systems often separate map exploration and path…

Robotics · Computer Science 2025-07-24 Yuxuan Zhang , Adnan Abdullah , Sanjeev J. Koppal , Md Jahidul Islam

Dynamic scenes contain intricate spatio-temporal information, crucial for mobile robots, UAVs, and autonomous driving systems to make informed decisions. Parsing these scenes into semantic triplets <Subject-Predicate-Object> for accurate…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Hang Zhang , Zhuoling Li , Jun Liu

Vision-Language Models (VLMs) are increasingly applied in autonomous driving for unified perception and reasoning, but high inference latency hinders real-time deployment. Early-exit reduces latency by terminating inference at intermediate…

Robotics · Computer Science 2025-10-13 Haibo Hu , Lianming Huang , Xinyu Wang , Yufei Cui , Shangyu Wu , Nan Guan , Chun Jason Xue

Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. However, existing SNNs still exhibit a significant performance gap compared to Artificial Neural Networks (ANNs) due to inadequate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Xuerui Qiu , Peixi Wu , Yaozhi Wen , Shaowei Gu , Yuqi Pan , Xinhao Luo , Bo XU , Guoqi Li

Language models (LMs) are increasingly applied to robotic navigation; however, existing benchmarks primarily emphasize navigation success rates while paying limited attention to social compliance. Moreover, relying on large-scale LMs can…

Robotics · Computer Science 2026-03-24 Ling Xiao , Daeun Song , Xuesu Xiao , Toshihiko Yamasaki

We present an optimization study of the Vision-Language Frontier Maps (VLFM) applied to the Object Goal Navigation task in robotics. Our work evaluates the efficiency and performance of various vision-language models, object detectors,…

Robotics · Computer Science 2025-07-03 Dmytro Kuzmenko , Nadiya Shvai
‹ Prev 1 4 5 6 7 8 10 Next ›