中文
相关论文

相关论文: Where Are You? Localization from Embodied Dialog

200 篇论文

We address the challenging task of Localization via Embodied Dialog (LED). Given a dialog from two agents, an Observer navigating through an unknown environment and a Locator who is attempting to identify the Observer's location, the goal…

计算机视觉与模式识别 · 计算机科学 2022-10-11 Meera Hahn , James M. Rehg

Multimodal learning has advanced the performance for many vision-language tasks. However, most existing works in embodied dialog research focus on navigation and leave the localization task understudied. The few existing dialog-based…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Chao Zhang , Mohan Li , Ignas Budvytis , Stephan Liwicki

Robots navigating in human environments should use language to ask for assistance and be able to understand human responses. To study this challenge, we introduce Cooperative Vision-and-Dialog Navigation, a dataset of over 2k embodied,…

计算与语言 · 计算机科学 2019-10-15 Jesse Thomason , Michael Murray , Maya Cakmak , Luke Zettlemoyer

In this work, we study the problem of Embodied Referring Expression Grounding, where an agent needs to navigate in a previously unseen environment and localize a remote object described by a concise high-level natural language instruction.…

计算机视觉与模式识别 · 计算机科学 2022-12-05 Mingxiao Li , Zehao Wang , Tinne Tuytelaars , Marie-Francine Moens

Predicting where people can walk in a scene is important for many tasks, including autonomous driving systems and human behavior analysis. Yet learning a computational model for this purpose is challenging due to semantic ambiguity and a…

计算机视觉与模式识别 · 计算机科学 2020-08-21 Jin Sun , Hadar Averbuch-Elor , Qianqian Wang , Noah Snavely

A crucial ability of mobile intelligent agents is to integrate the evidence from multiple sensory inputs in an environment and to make a sequence of actions to reach their goals. In this paper, we attempt to approach the problem of…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Chuang Gan , Yiwei Zhang , Jiajun Wu , Boqing Gong , Joshua B. Tenenbaum

Embodied intelligence fundamentally requires a capability to determine where to act in 3D space. We formalize this requirement as embodied localization -- the problem of predicting executable 3D points conditioned on visual observations and…

机器人学 · 计算机科学 2026-03-31 Qiming Zhu , Zhirui Fang , Tianming Zhang , Chuanxiu Liu , Xiaoke Jiang , Lei Zhang

Situated embodied conversation requires robots to interleave real-time dialogue with active perception: deciding what to look at, when to look, and what to say under tight latency constraints. We present a simple, minimal system recipe that…

机器人学 · 计算机科学 2026-02-05 Dong Won Lee , Sarah Gillet , Louis-Philippe Morency , Cynthia Breazeal , Hae Won Park

While current visual captioning models have achieved impressive performance, they often assume that the image is well-captured and provides a complete view of the scene. In real-world scenarios, however, a single image may not offer a good…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Anwen Hu , Shizhe Chen , Liang Zhang , Qin Jin

Effective collaboration between embodied agents requires more than acting in a shared environment; it demands communication grounded in each agent's evolving understanding of the world. When agents can only partially observe their…

多智能体系统 · 计算机科学 2026-05-19 Vardhan Dongre , Dilek Hakkani-Tür

We study the problem of jointly reasoning about language and vision through a navigation and spatial reasoning task. We introduce the Touchdown task and dataset, where an agent must first follow navigation instructions in a real-life visual…

计算机视觉与模式识别 · 计算机科学 2020-05-19 Howard Chen , Alane Suhr , Dipendra Misra , Noah Snavely , Yoav Artzi

There has been a significant recent progress in the field of Embodied AI with researchers developing models and algorithms enabling embodied agents to navigate and interact within completely unseen environments. In this paper, we propose a…

计算机视觉与模式识别 · 计算机科学 2021-03-31 Luca Weihs , Matt Deitke , Aniruddha Kembhavi , Roozbeh Mottaghi

Recent advances in deep thinking models have demonstrated remarkable reasoning capabilities on mathematical and coding tasks. However, their effectiveness in embodied domains which require continuous interaction with environments through…

Embodied instruction following is a challenging problem requiring an agent to infer a sequence of primitive actions to achieve a goal environment state from complex language and visual inputs. Action Learning From Realistic Environments and…

人工智能 · 计算机科学 2021-01-12 Shane Storks , Qiaozi Gao , Govind Thattai , Gokhan Tur

We introduce DialNav, a novel collaborative embodied dialog task, where a navigation agent (Navigator) and a remote guide (Guide) engage in multi-turn dialog to reach a goal location. Unlike prior work, DialNav aims for holistic evaluation…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Leekyeung Han , Hyunji Min , Gyeom Hwangbo , Jonghyun Choi , Paul Hongsuck Seo

In the realm of computer vision and robotics, embodied agents are expected to explore their environment and carry out human instructions. This necessitates the ability to fully understand 3D scenes given their first-person observations and…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Tai Wang , Xiaohan Mao , Chenming Zhu , Runsen Xu , Ruiyuan Lyu , Peisen Li , Xiao Chen , Wenwei Zhang , Kai Chen , Tianfan Xue , Xihui Liu , Cewu Lu , Dahua Lin , Jiangmiao Pang

Visual grounding in 3D is the key for embodied agents to localize language-referred objects in open-world environments. However, existing benchmarks are limited to indoor focus, single-platform constraints, and small scale. We introduce…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Rong Li , Yuhao Dong , Tianshuai Hu , Ao Liang , Youquan Liu , Dongyue Lu , Liang Pan , Lingdong Kong , Junwei Liang , Ziwei Liu

Autonomous robot systems for applications from search and rescue to assistive guidance should be able to engage in natural language dialog with people. To study such cooperative communication, we introduce Robot Simultaneous Localization…

机器人学 · 计算机科学 2020-10-27 Shurjo Banerjee , Jesse Thomason , Jason J. Corso

Nowadays, users are encouraged to activate across multiple online social networks simultaneously. Anchor link prediction, which aims to reveal the correspondence among different accounts of the same user across networks, has been regarded…

社会与信息网络 · 计算机科学 2021-08-26 Jiangli Shao , Yongqing Wang , Hao Gao , Huawei Shen , Yangyang Li , Xueqi Cheng

Robots operating in human spaces must be able to engage in natural language interaction with people, both understanding and executing instructions, and using conversation to resolve ambiguity and recover from mistakes. To study this, we…

‹ 上一页 1 2 3 10 下一页 ›