中文
相关论文

相关论文: Can an Embodied Agent Find Your "Cat-shaped Mug"? …

200 篇论文

We present an active mapping system that plans for both long-horizon exploration goals and short-term actions using a 3D Gaussian Splatting (3DGS) representation. Existing methods either do not take advantage of recent developments in…

机器人学 · 计算机科学 2025-09-08 Wen Jiang , Boshu Lei , Katrina Ashton , Kostas Daniilidis

Benefiting from language flexibility and compositionality, humans naturally intend to use language to command an embodied agent for complex tasks such as navigation and object manipulation. In this work, we aim to fill the blank of the last…

机器人学 · 计算机科学 2022-08-18 Kaizhi Zheng , Xiaotong Chen , Odest Chadwicke Jenkins , Xin Eric Wang

Despite advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), their integration into language-grounded, human-like embodied agents remains incomplete, hindering complex real-life task performance in physical…

计算与语言 · 计算机科学 2024-08-20 Zhili Cheng , Zhitong Wang , Jinyi Hu , Shengding Hu , An Liu , Yuge Tu , Pengkai Li , Lei Shi , Zhiyuan Liu , Maosong Sun

Capitalizing on the remarkable advancements in Large Language Models (LLMs), there is a burgeoning initiative to harness LLMs for instruction following robotic navigation. Such a trend underscores the potential of LLMs to generalize…

计算机视觉与模式识别 · 计算机科学 2024-09-23 Gengze Zhou , Yicong Hong , Zun Wang , Xin Eric Wang , Qi Wu

In recent years, the rapid advancement of Large Language Models (LLMs) such as the Generative Pre-trained Transformer (GPT) has attracted increasing attention due to their potential in a variety of practical applications. The application of…

人工智能 · 计算机科学 2025-04-24 Jinzhou Lin , Han Gao , Xuxiang Feng , Rongtao Xu , Changwei Wang , Man Zhang , Li Guo , Shibiao Xu

Zero-Shot Object Navigation (ZSON) requires agents to autonomously locate and approach unseen objects in unfamiliar environments and has emerged as a particularly challenging task within the domain of Embodied AI. Existing datasets for…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Ji Ma , Hongming Dai , Yao Mu , Pengying Wu , Hao Wang , Xiaowei Chi , Yang Fei , Shanghang Zhang , Chang Liu

Recent approaches have shown that training deep neural networks directly on large-scale image-text pair collections enables zero-shot transfer on various recognition tasks. One central issue is how this can be generalized to object…

计算机视觉与模式识别 · 计算机科学 2022-08-30 Johnathan Xie , Shuai Zheng

Vision-and-Language Navigation (VLN) has recently benefited from Multimodal Large Language Models (MLLMs), enabling zero-shot navigation. While recent exploration-based zero-shot methods have shown promising results by leveraging global…

机器人学 · 计算机科学 2026-03-31 Jiwen Zhang , Xiangyu Shi , Siyuan Wang , Zerui Li , Zhongyu Wei , Qi Wu

Although learning-based vision-and-language navigation (VLN) agents can learn spatial knowledge implicitly from large-scale training data, zero-shot VLN agents lack this process, relying primarily on local observations for navigation, which…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Jiwen Zhang , Zejun Li , Siyuan Wang , Xiangyu Shi , Zhongyu Wei , Qi Wu

Reasoning about spatial relationships between objects is essential for many real-world robotic tasks, such as fetch-and-delivery, object rearrangement, and object search. The ability to detect and disambiguate different objects and identify…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Negar Nejatishahidin , Madhukar Reddy Vongala , Jana Kosecka

We study zero-shot instance navigation, in which the agent navigates to a specific object without using object annotations for training. Previous object navigation approaches apply the image-goal navigation (ImageNav) task (go to the…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Xinyu Sun , Lizhao Liu , Hongyan Zhi , Ronghe Qiu , Junwei Liang

Embodied scene understanding serves as the cornerstone for autonomous agents to perceive, interpret, and respond to open driving scenarios. Such understanding is typically founded upon Vision-Language Models (VLMs). Nevertheless, existing…

计算机视觉与模式识别 · 计算机科学 2024-03-08 Yunsong Zhou , Linyan Huang , Qingwen Bu , Jia Zeng , Tianyu Li , Hang Qiu , Hongzi Zhu , Minyi Guo , Yu Qiao , Hongyang Li

Embodied long-horizon manipulation requires robotic systems to process multimodal inputs-such as vision and natural language-and translate them into executable actions. However, existing learning-based approaches often depend on large,…

机器人学 · 计算机科学 2025-08-22 Yuan Meng , Xiangtong Yao , Haihui Ye , Yirui Zhou , Shengqiang Zhang , Zhenguo Sun , Xukun Li , Zhenshan Bing , Alois Knoll

Large language models (LLMs) have undergone significant expansion and have been increasingly integrated across various domains. Notably, in the realm of robot task planning, LLMs harness their advanced reasoning and language comprehension…

Zero-shot open-vocabulary object navigation has progressed rapidly with the emergence of large Vision-Language Models (VLMs) and Large Language Models (LLMs), now widely used as high-level decision-makers instead of end-to-end policies.…

Next location prediction plays a crucial role in various real-world applications. Recently, due to the limitation of existing deep learning methods, attempts have been made to apply large language models (LLMs) to zero-shot next location…

机器学习 · 计算机科学 2025-02-11 Jie Feng , Yuwei Du , Jie Zhao , Yong Li

In recent years, as machine learning, particularly for vision and language understanding, has been improved, research in embedded AI has also evolved. VOYAGER is a well-known LLM-based embodied AI that enables autonomous exploration in the…

人工智能 · 计算机科学 2024-06-05 Wakana Haijima , Kou Nakakubo , Masahiro Suzuki , Yutaka Matsuo

Vision-and-Language Navigation (VLN), as a crucial research problem of Embodied AI, requires an embodied agent to navigate through complex 3D environments following natural language instructions. Recent research has highlighted the…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Bingqian Lin , Yunshuang Nie , Ziming Wei , Jiaqi Chen , Shikui Ma , Jianhua Han , Hang Xu , Xiaojun Chang , Xiaodan Liang

A robot in a human-centric environment needs to account for the human's intent and future motion in its task and motion planning to ensure safe and effective operation. This requires symbolic reasoning about probable future actions and the…

机器人学 · 计算机科学 2023-11-01 Moritz A. Graule , Volkan Isler

In this paper, we focus on the problem of efficiently locating a target object described with free-form language using a mobile robot equipped with vision sensors (e.g., an RGBD camera). Conventional active visual search predefines a set of…

机器人学 · 计算机科学 2023-02-08 Jeongeun Park , Taerim Yoon , Jejoon Hong , Youngjae Yu , Matthew Pan , Sungjoon Choi