中文
相关论文

相关论文: Learning Semantic-Agnostic and Spatial-Aware Repre…

200 篇论文

We are concerned with the question of how an agent can acquire its own representations from sensory data. We restrict our focus to learning representations for long-term planning, a class of problems that state-of-the-art learning methods…

机器学习 · 计算机科学 2022-05-05 Steven James , Benjamin Rosman , George Konidaris

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location in 3D environments following the natural language instruction. In this field, the agent is usually trained and evaluated in the navigation simulators,…

机器人学 · 计算机科学 2024-10-15 Zihan Wang , Xiangyang Li , Jiahao Yang , Yeqi Liu , Shuqiang Jiang

We introduce environment predictive coding, a self-supervised approach to learn environment-level representations for embodied agents. In contrast to prior work on self-supervised learning for images, we aim to jointly encode a series of…

计算机视觉与模式识别 · 计算机科学 2021-02-05 Santhosh K. Ramakrishnan , Tushar Nagarajan , Ziad Al-Halah , Kristen Grauman

Object goal navigation (ObjectNav) in unseen environments is a fundamental task for Embodied AI. Agents in existing works learn ObjectNav policies based on 2D maps, scene graphs, or image sequences. Considering this task happens in 3D…

机器人学 · 计算机科学 2023-04-03 Jiazhao Zhang , Liu Dai , Fanpeng Meng , Qingnan Fan , Xuelin Chen , Kai Xu , He Wang

Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic information between the…

计算机视觉与模式识别 · 计算机科学 2020-06-15 Karren Yang , Bryan Russell , Justin Salamon

Open-vocabulary scene understanding is crucial for robotic applications, enabling robots to comprehend complex 3D environmental contexts and supporting various downstream tasks such as navigation and manipulation. However, existing methods…

机器人学 · 计算机科学 2026-03-19 Siting Zhu , Ziyun Lu , Guangming Wang , Chenguang Huang , Yongbo Chen , I-Ming Chen , Wolfram Burgard , Hesheng Wang

Autonomous agents, particularly in the field of robotics, rely on sensory information to perceive and navigate their environment. However, these sensory inputs are often imperfect, leading to distortions in the agent's internal…

机器人学 · 计算机科学 2025-07-11 David Warutumo , Ciira wa Maina

Sight and hearing are two senses that play a vital role in human communication and scene understanding. To mimic human perception ability, audio-visual learning, aimed at developing computational approaches to learn from both audio and…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Yake Wei , Di Hu , Yapeng Tian , Xuelong Li

Holistic scene understanding poses a fundamental contribution to the autonomous operation of a robotic agent in its environment. Key ingredients include a well-defined representation of the surroundings to capture its spatial structure as…

机器人学 · 计算机科学 2024-05-24 Niclas Vödisch

Object Goal Navigation (ObjectNav) task is to navigate an agent to an object category in unseen environments without a pre-built map. In this paper, we solve this task by predicting the distance to the target using semantically-related…

机器人学 · 计算机科学 2022-07-14 Minzhao Zhu , Binglei Zhao , Tao Kong

Intelligent embodied agents (e.g. robots) need to perform complex semantic tasks in unfamiliar environments. Among many skills that the agents need to possess, building and maintaining a semantic map of the environment is most crucial in…

机器人学 · 计算机科学 2025-08-13 Sonia Raychaudhuri , Angel X. Chang

We study open-world 3D scene understanding, a family of tasks that require agents to reason about their 3D environment with an open-set vocabulary and out-of-domain visual inputs - a critical skill for robots to operate in the unstructured…

计算机视觉与模式识别 · 计算机科学 2022-12-07 Huy Ha , Shuran Song

We study the task of zero-shot vision-and-language navigation (ZS-VLN), a practical yet challenging problem in which an agent learns to navigate following a path described by language instructions without requiring any path-instruction…

计算机视觉与模式识别 · 计算机科学 2023-08-17 Peihao Chen , Xinyu Sun , Hongyan Zhi , Runhao Zeng , Thomas H. Li , Gaowen Liu , Mingkui Tan , Chuang Gan

Vision-and-Language Navigation (VLN) is a multi-modal, cooperative task requiring agents to interpret human instructions, navigate 3D environments, and communicate effectively under ambiguity. This paper presents a comprehensive review of…

机器人学 · 计算机科学 2025-12-02 Nivedan Yakolli , Avinash Gautam , Abhijit Das , Yuankai Qi , Virendra Singh Shekhawat

Natural language instructions for visual navigation often use scene descriptions (e.g., "bedroom") and object references (e.g., "green chairs") to provide a breadcrumb trail to a goal location. This work presents a transformer-based…

计算机视觉与模式识别 · 计算机科学 2021-10-28 Abhinav Moudgil , Arjun Majumdar , Harsh Agrawal , Stefan Lee , Dhruv Batra

The sound of crashing waves, the roar of fast-moving cars -- sound conveys important information about the objects in our surroundings. In this work, we show that ambient sounds can be used as a supervisory signal for learning visual…

计算机视觉与模式识别 · 计算机科学 2017-12-21 Andrew Owens , Jiajun Wu , Josh H. McDermott , William T. Freeman , Antonio Torralba

Natural human-robot interaction in complex and unpredictable environments is one of the main research lines in robotics. In typical real-world scenarios, humans are at some distance from the robot and the acquired signals are strongly…

机器人学 · 计算机科学 2015-09-04 Xavier Alameda-Pineda , Radu Horaud

Map representations learned by expert demonstrations have shown promising research value. However, the field of visual navigation still faces challenges due to the lack of real-world human-navigation datasets that can support efficient,…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Faith Johnson , Bryan Bo Cao , Kristin Dana , Shubham Jain , Ashwin Ashok

Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit visual cues as a…

声音 · 计算机科学 2025-03-18 Swapnil Bhosale , Haosen Yang , Diptesh Kanojia , Jiankang Deng , Xiatian Zhu

Semantic navigation is necessary to deploy mobile robots in uncontrolled environments like our homes, schools, and hospitals. Many learning-based approaches have been proposed in response to the lack of semantic understanding of the…

机器人学 · 计算机科学 2022-12-05 Theophile Gervet , Soumith Chintala , Dhruv Batra , Jitendra Malik , Devendra Singh Chaplot