中文
相关论文

相关论文: TINA: Think, Interaction, and Action Framework for…

200 篇论文

Vision-and-Language Navigation (VLN) aims to enable an embodied agent to follow natural-language instructions and navigate to a target location in unseen 3D environments. We argue that adapting VLMs to VLN requires endowing them with two…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Pengna Li , Kangyi Wu , Shaoqing Xu , Fang Li , Hanbing Li , Lin Zhao , Kailin Lyu , Long Chen , Zhi-Xin Yang , Nanning Zheng

Recent advances in large vision-language models (VLMs) and large language models (LLMs) have enabled zero-shot approaches to visual language navigation (VLN), where an agent follows natural language instructions using only ego perception…

机器人学 · 计算机科学 2026-03-03 Chaoran Xiong , Litao Wei , Xinhao Hu , Kehui Ma , Ziyi Xia , Zixin Jiang , Zhen Sun , Ling Pei

Object Goal Navigation (ObjectNav) challenges robots to find objects in unseen environments, demanding sophisticated reasoning. While Vision-Language Models (VLMs) show potential, current ObjectNav methods often employ them superficially,…

机器人学 · 计算机科学 2025-06-23 Mobin Habibpour , Fatemeh Afghah

Vision-and-language navigation (VLN) is a multimodal task where an agent follows natural language instructions and navigates in visual environments. Multiple setups have been proposed, and researchers apply new model architectures or…

计算机视觉与模式识别 · 计算机科学 2022-05-05 Wanrong Zhu , Yuankai Qi , Pradyumna Narayana , Kazoo Sone , Sugato Basu , Xin Eric Wang , Qi Wu , Miguel Eckstein , William Yang Wang

Rapid adaptation in unseen environments is essential for scalable real-world autonomy, yet existing approaches rely on exhaustive exploration or rigid navigation policies that fail to generalize. We present VLN-Zero, a two-phase…

机器人学 · 计算机科学 2025-09-24 Neel P. Bhatt , Yunhao Yang , Rohan Siva , Pranay Samineni , Daniel Milan , Zhangyang Wang , Ufuk Topcu

LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) tasks. However, most zero-shot methods primarily rely on closed-source LLMs as navigators, which face challenges related to high token…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Kailin Lyu , Kangyi Wu , Pengna Li , Xiuyu Hu , Qingyi Si , Cui Miao , Ning Yang , Zihang Wang , Long Xiao , Lianyu Hu , Jingyuan Sun , Ce Hao

Visual navigation is an essential skill for home-assistance robots, providing the object-searching ability to accomplish long-horizon daily tasks. Many recent approaches use Large Language Models (LLMs) for commonsense inference to improve…

机器人学 · 计算机科学 2024-10-15 Xinxin Zhao , Wenzhe Cai , Likun Tang , Teng Wang

This paper considers a scenario in city navigation: an AI agent is provided with language descriptions of the goal location with respect to some well-known landmarks; By only observing the scene around, including recognizing landmarks and…

人工智能 · 计算机科学 2024-10-18 Qingbin Zeng , Qinglong Yang , Shunan Dong , Heming Du , Liang Zheng , Fengli Xu , Yong Li

Autonomous navigation under natural language instructions represents a crucial step toward embodied intelligence, enabling complex task execution in environments ranging from industrial facilities to domestic spaces. However,…

机器人学 · 计算机科学 2026-03-05 Hongyu Song , Rishabh Dev Yadav , Cheng Guo , Wei Pan

Vision-and-Language Navigation (VLN) poses significant challenges for agents to interpret natural language instructions and navigate complex 3D environments. While recent progress has been driven by large-scale pre-training and data…

人工智能 · 计算机科学 2026-05-14 Tianyi Ma , Yue Zhang , Zehao Wang , Parisa Kordjamshidi

In recent years, the field of indoor navigation has witnessed groundbreaking advancements through the integration of Large Language Models (LLMs). Traditional navigation approaches relying on pre-built maps or reinforcement learning exhibit…

机器人学 · 计算机科学 2025-04-23 Anlong Zhang , Jianmin Ji

Vision-and-Language Navigation (VLN) requires an embodied agent to navigate in a complex 3D environment according to natural language instructions. Recent progress in large language models (LLMs) has enabled language-driven navigation with…

机器人学 · 计算机科学 2026-01-27 Zijun Li , Shijie Li , Zhenxi Zhang , Bin Li , Shoujun Zhou

Methods for navigation based on large-scale learning typically treat each episode as a new problem, where the agent is spawned with a clean memory in an unknown environment. While these generalization capabilities to an unknown environment…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Gianluca Monaci , Rafael S. Rezende , Romain Deffayet , Gabriela Csurka , Guillaume Bono , Hervé Déjean , Stéphane Clinchant , Christian Wolf

Zero-Shot Object Navigation (ZSON) enables agents to navigate towards open-vocabulary objects in unknown environments. The existing works of ZSON mainly focus on following individual instructions to find generic object classes, neglecting…

机器人学 · 计算机科学 2024-05-31 Yinpei Dai , Run Peng , Sikai Li , Joyce Chai

Breakthrough progress in vision-based navigation through unknown environments has been achieved by using multimodal large language models (MLLMs). These models can plan a sequence of motions by evaluating the current view at each time step…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Wanrong Zheng , Yunhao Ge , Laurent Itti

Vision-and-Language Navigation (VLN) presents a complex challenge in embodied AI, requiring agents to interpret natural language instructions and navigate through visually rich, unfamiliar environments. Recent advances in large…

机器人学 · 计算机科学 2025-06-13 Yicheng Duan , Kaiyu tang

Vision-language navigation (VLN) is the task of entailing an agent to carry out navigational instructions inside photo-realistic environments. One of the key challenges in VLN is how to conduct a robust navigation by mitigating the…

计算机视觉与模式识别 · 计算机科学 2020-08-21 Hanqing Wang , Wenguan Wang , Tianmin Shu , Wei Liang , Jianbing Shen

Vision-and-Language Navigation (VLN) requires an agent to interpret natural language instructions and navigate complex environments. Current approaches often adopt a "black-box" paradigm, where a single Large Language Model (LLM) makes…

计算机视觉与模式识别 · 计算机科学 2025-09-11 Chenghao Liu , Zhimu Zhou , Jiachen Zhang , Minghao Zhang , Songfang Huang , Huiling Duan

Recent advances in Vision-and-Language Navigation in Continuous Environments (VLN-CE) have leveraged multimodal large language models (MLLMs) to achieve zero-shot navigation. However, existing methods often rely on panoramic observations…

机器人学 · 计算机科学 2025-11-04 Xiangyu Shi , Zerui Li , Yanyuan Qiao , Qi Wu

Zero-shot object navigation requires agents to locate unseen target objects in unfamiliar environments without prior maps or task-specific training which remains a significant challenge. Although recent advancements in vision-language…

机器人学 · 计算机科学 2026-04-08 Feng Wu , Wei Zuo , Wenliang Yang , Jun Xiao , Yang Liu , Xinhua Zeng