English
Related papers

Related papers: Learning a Visually Grounded Memory Assistant

200 papers

This work presents a modular architecture for simultaneous mapping and target driven navigation in indoors environments. The semantic and appearance stored in 2.5D map is distilled from RGB images, semantic segmentation and outputs of…

Computer Vision and Pattern Recognition · Computer Science 2019-11-20 Georgios Georgakis , Yimeng Li , Jana Kosecka

There has been a significant recent progress in the field of Embodied AI with researchers developing models and algorithms enabling embodied agents to navigate and interact within completely unseen environments. In this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Luca Weihs , Matt Deitke , Aniruddha Kembhavi , Roozbeh Mottaghi

One application area of long-term memory (LTM) capabilities with increasing traction is personal AI companions and assistants. With the ability to retain and contextualize past interactions and adapt to user preferences, personal AI…

Computers and Society · Computer Science 2024-09-18 Eunhae Lee

Human intelligence's adaptability is remarkable, allowing us to adjust to new tasks and multi-modal environments swiftly. This skill is evident from a young age as we acquire new abilities and solve problems by imitating others or following…

Artificial Intelligence · Computer Science 2023-05-19 Shrestha Mohanty , Negar Arabzadeh , Julia Kiseleva , Artem Zholus , Milagro Teruel , Ahmed Awadallah , Yuxuan Sun , Kavya Srinet , Arthur Szlam

Adaptive interfaces can help users perform sequential decision-making tasks like robotic teleoperation given noisy, high-dimensional command signals (e.g., from a brain-computer interface). Recent advances in human-in-the-loop machine…

Robotics · Computer Science 2023-09-08 Jensen Gao , Siddharth Reddy , Glen Berseth , Anca D. Dragan , Sergey Levine

We investigated the human capacity to acquire multiple visuomotor mappings for de novo skills. Using a grid navigation paradigm, we tested whether contextual cues implemented as different "grid worlds", allow participants to learn two…

Neurons and Cognition · Quantitative Biology 2024-02-06 Carlos A. Velazquez-Vargas , Isaac Ray Christian , Jordan A. Taylor , Sreejan Kumar

Robots are often required to operate in environments where humans are not present, but yet require the human context information for better human-robot interaction. Even when humans are present in the environment, detecting their presence…

Computer Vision and Pattern Recognition · Computer Science 2019-06-14 Lasitha Piyathilaka , Sarath Kodagoda

Artificial learning systems aspire to mimic human intelligence by continually learning from a stream of tasks without forgetting past knowledge. One way to enable such learning is to store past experiences in the form of input examples in…

Machine Learning · Computer Science 2022-10-13 Gobinda Saha , Kaushik Roy

In image-based camera localization systems, information about the environment is usually stored in some representation, which can be referred to as a map. Conventionally, most maps are built upon hand-crafted features. Recently, neural…

Computer Vision and Pattern Recognition · Computer Science 2019-04-17 Mingpan Guo , Stefan Matthes , Jiaojiao Ye , Hao Shen

As a fundamental problem for Artificial Intelligence, multi-agent system (MAS) is making rapid progress, mainly driven by multi-agent reinforcement learning (MARL) techniques. However, previous MARL methods largely focused on grid-world…

Computer Vision and Pattern Recognition · Computer Science 2021-07-21 Haiyang Wang , Wenguan Wang , Xizhou Zhu , Jifeng Dai , Liwei Wang

Using touch devices to navigate in virtual 3D environments such as computer assisted design (CAD) models or geographical information systems (GIS) is inherently difficult for humans, as the 3D operations have to be performed by the user on…

Machine Learning · Computer Science 2019-08-29 Quentin Debard , Jilles Steeve Dibangoye , Stéphane Canu , Christian Wolf

Visualization plays a relevant role for discovering patterns in big sets of data. In fact, the most common way to help a human with a pattern interpretation is through a graphic. In 2D/3D virtual environments for procedural training the…

Human-Computer Interaction · Computer Science 2025-12-24 Diego Riofrío-Luzcando , Jaime RamÍrez , Cristian Moral , Angélica de Antonio , Marta Berrocal-Lobo

We present an optimised multi-modal dialogue agent for interactive learning of visually grounded word meanings from a human tutor, trained on real human-human tutoring data. Within a life-long interactive learning period, the agent, trained…

Computation and Language · Computer Science 2017-10-02 Yanchao Yu , Arash Eshghi , Oliver Lemon

Grounding language to the visual observations of a navigating agent can be performed using off-the-shelf visual-language models pretrained on Internet-scale data (e.g., image captions). While this is useful for matching images to natural…

Robotics · Computer Science 2023-03-09 Chenguang Huang , Oier Mees , Andy Zeng , Wolfram Burgard

In this study, we propose a novel human-like memory architecture designed for enhancing the cognitive abilities of large language model based dialogue agents. Our proposed architecture enables agents to autonomously recall memories…

Human-Computer Interaction · Computer Science 2024-04-02 Yuki Hou , Haruki Tamoto , Homei Miyashita

Strategies for finding one's way through an unfamiliar environment may be helped by computer generated 2D maps, 3D virtual environments, or other navigation aids. The relative effectiveness of 2D and 3D virtual navigation aids was…

Human-Computer Interaction · Computer Science 2018-05-24 Yasmine Boumenir , Fanny Georges , Jeremy Valentin , Guy Rebillard , Birgitta Dresp-Langley

Image-goal navigation is a challenging task that requires an agent to navigate to a goal indicated by an image in unfamiliar environments. Existing methods utilizing diverse scene memories suffer from inefficient exploration since they use…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Hongxin Li , Zeyu Wang , Xu Yang , Yuran Yang , Shuqi Mei , Zhaoxiang Zhang

Visual navigation is an essential skill for home-assistance robots, providing the object-searching ability to accomplish long-horizon daily tasks. Many recent approaches use Large Language Models (LLMs) for commonsense inference to improve…

Robotics · Computer Science 2024-10-15 Xinxin Zhao , Wenzhe Cai , Likun Tang , Teng Wang

When humans perform a task, such as playing a game, they selectively pay attention to certain parts of the visual input, gathering relevant information and sequentially combining it to build a representation from the sensory data. In this…

Artificial Intelligence · Computer Science 2018-07-26 Khimya Khetarpal , Doina Precup

Vision-and-language navigation (VLN) agents are trained to navigate in real-world environments by following natural language instructions. A major challenge in VLN is the limited availability of training data, which hinders the models'…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Zi-Yi Dou , Feng Gao , Nanyun Peng