English
Related papers

Related papers: Endowing Embodied Agents with Spatial Reasoning Ca…

200 papers

Recent advances in vision-language models have made zero-shot navigation feasible, enabling robots to follow natural language instructions without requiring labeling. However, existing methods that explicitly store language vectors in grid…

Robotics · Computer Science 2026-02-13 Sibaek Lee , Hyeonwoo Yu , Giseop Kim , Sunwook Choi

Grounding open-ended semantic instructions into physically executable local goals is a fundamental challenge in human-robot interaction. While existing navigation frameworks often regress deterministic waypoints, this rigid formulation…

Robotics · Computer Science 2026-05-20 Kaijie Yun , Yue Chen

Vision-Language Navigation (VLN) requires an embodied agent to navigate complex environments by following natural language instructions, which typically demands tight fusion of visual and language modalities. Existing VLN methods often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Daojie Peng , Fulong Ma , Jun Ma

Embodied navigation has long been fragmented by task-specific architectures. We introduce ABot-N0, a unified Vision-Language-Action (VLA) foundation model that achieves a ``Grand Unification'' across 5 core tasks: Point-Goal, Object-Goal,…

Vision-Language Navigation (VLN) enables robots to follow natural-language instructions in visually grounded environments, serving as a key capability for embodied robotic systems. Recent Vision-Language-Action (VLA) models have…

Robotics · Computer Science 2026-03-09 Qitong Wang , Yijun Liang , Ming Li , Tianyi Zhou , Christopher Rasmussen

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jinzhou Tang , Jusheng zhang , Sidi Liu , Waikit Xiu , Qinhan Lv , Xiying Li

Vision-and-Language Navigation (VLN) presents a complex challenge in embodied AI, requiring agents to interpret natural language instructions and navigate through visually rich, unfamiliar environments. Recent advances in large…

Robotics · Computer Science 2025-06-13 Yicheng Duan , Kaiyu tang

Building on the unprecedented capabilities of large language models for command understanding and zero-shot recognition of multi-modal vision-language transformers, visual language navigation (VLN) has emerged as an effective way to address…

Robotics · Computer Science 2024-07-11 Chashi Mahiul Islam , Shaeke Salman , Montasir Shams , Xiuwen Liu , Piyush Kumar

Vision-and-language navigation (VLN) stands as a key research problem of Embodied AI, aiming at enabling agents to navigate in unseen environments following linguistic instructions. In this field, generalization is a long-standing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Jiazhao Zhang , Kunyu Wang , Rongtao Xu , Gengze Zhou , Yicong Hong , Xiaomeng Fang , Qi Wu , Zhizheng Zhang , He Wang

Autonomous navigation guided by natural language instructions in embodied environments remains a challenge for vision-language navigation (VLN) agents. Although recent advancements in learning diverse and fine-grained visual environmental…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Xuesong Zhang , Jia Li , Yunbo Xu , Zhenzhen Hu , Richang Hong

We introduce HY-Embodied-0.5, a family of foundation models specifically designed for real-world embodied agents. To bridge the gap between general Vision-Language Models (VLMs) and the demands of embodied agents, our models are developed…

Vision-and-Language Navigation (VLN) tasks require an agent to follow textual instructions to navigate through 3D environments. Traditional approaches use supervised learning methods, relying heavily on domain-specific datasets to train VLN…

Robotics · Computer Science 2025-02-12 Yanyuan Qiao , Wenqi Lyu , Hui Wang , Zixu Wang , Zerui Li , Yuan Zhang , Mingkui Tan , Qi Wu

Real world visual navigation requires robots to operate in unfamiliar, human-occupied dynamic environments. Navigation around humans is especially difficult because it requires anticipating their future motion, which can be quite…

Robotics · Computer Science 2021-02-16 Varun Tolani , Somil Bansal , Aleksandra Faust , Claire Tomlin

As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere reachability to social compliance. However, current agents suffer from a "goal-driven trap",…

Artificial Intelligence · Computer Science 2026-04-21 Jiawen Wen , Penglei Sun , Wenjie Zhang , Suixuan Qiu , Weisheng Xu , Xiaofei Yang , Xiaowen Chu

Existing navigation methods are primarily designed for specific robot embodiments, limiting their generalizability across diverse robot platforms. In this paper, we introduce X-Nav, a novel framework for end-to-end cross-embodiment…

Robotics · Computer Science 2025-11-27 Haitong Wang , Aaron Hao Tan , Angus Fung , Goldie Nejat

Existing vision-and-language navigation (VLN) models primarily reason over past and current visual observations, while largely ignoring the future visual dynamics induced by actions. As a result, they often lack an effective understanding…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Haihong Hao , Lei Chen , Mingfei Han , Changlin Li , Dong An , Yuqiang Yang , Zhihui Li , Xiaojun Chang

Progress in Embodied AI has made it possible for end-to-end-trained agents to navigate in photo-realistic environments with high-level reasoning and zero-shot or language-conditioned behavior, but benchmarks are still dominated by…

Understanding spatial and visual information is essential for a navigation agent who follows natural language instructions. The current Transformer-based VLN agents entangle the orientation and vision information, which limits the gain from…

Computer Vision and Pattern Recognition · Computer Science 2022-09-27 Yue Zhang , Parisa Kordjamshidi

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Ziyu Zhu , Xilin Wang , Yixuan Li , Zhuofan Zhang , Xiaojian Ma , Yixin Chen , Baoxiong Jia , Wei Liang , Qian Yu , Zhidong Deng , Siyuan Huang , Qing Li

Human-robot collaboration, in which the robot intelligently assists the human with the upcoming task, is an appealing objective. To achieve this goal, the agent needs to be equipped with a fundamental collaborative navigation ability, where…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Changhao Li , Xinyu Sun , Peihao Chen , Jugang Fan , Zixu Wang , Yanxia Liu , Jinhui Zhu , Chuang Gan , Mingkui Tan