English
Related papers

Related papers: GOAT-Bench: A Benchmark for Multi-Modal Lifelong N…

200 papers

In the context of visual navigation, the capacity to map a novel environment is necessary for an agent to exploit its observation history in the considered place and efficiently reach known goals. This ability can be associated with spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-04-26 Pierre Marza , Laetitia Matignon , Olivier Simonin , Christian Wolf

Goal-conditioned navigation models for ground robots trained using supervised learning show promising zero-shot transfer, but their collision-avoidance capability nevertheless degrades under distribution shift, i.e. environmental, robot or…

Robotics · Computer Science 2026-04-22 Louis Dezons , Quentin Picard , Rémi Marsal , François Goulette , David Filliat

Empowering embodied agents, such as robots, with Artificial Intelligence (AI) has become increasingly important in recent years. A major challenge is task open-endedness. In practice, robots often need to perform tasks with novel goals that…

Artificial Intelligence · Computer Science 2023-12-12 William Wei Wang , Dongqi Han , Xufang Luo , Yifei Shen , Charles Ling , Boyu Wang , Dongsheng Li

Entity state tracking is a necessary component of world modeling that requires maintaining coherent representations of entities over time. Previous work has benchmarked entity tracking performance in purely text-based tasks. We introduce…

Computation and Language · Computer Science 2026-02-10 Vanya Cohen , Raymond Mooney

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, which requires multi-round dialogue spatial reasoning and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Xunyi Zhao , Gengze Zhou , Qi Wu

Audio-visual embodied navigation, as a hot research topic, aims training a robot to reach an audio target using egocentric visual (from the sensors mounted on the robot) and audio (emitted from the target) input. The audio-visual…

Sound · Computer Science 2022-10-06 Yinfeng Yu , Lele Cao , Fuchun Sun , Xiaohong Liu , Liejun Wang

Multimodal retrieval is becoming a crucial component of modern AI applications, yet its evaluation lags behind the demands of more realistic and challenging scenarios. Existing benchmarks primarily probe surface-level semantic…

Information Retrieval · Computer Science 2025-10-01 Junjie Zhou , Ze Liu , Lei Xiong , Jin-Ge Yao , Yueze Wang , Shitao Xiao , Fenfen Lin , Miguel Hu Chen , Zhicheng Dou , Siqi Bao , Defu Lian , Yongping Xiong , Zheng Liu

Embodied AI has been recently gaining attention as it aims to foster the development of autonomous and intelligent agents. In this paper, we devise a novel embodied setting in which an agent needs to explore a previously unknown environment…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Roberto Bigazzi , Federico Landi , Marcella Cornia , Silvia Cascianelli , Lorenzo Baraldi , Rita Cucchiara

Image-goal navigation is a challenging task, as it requires the agent to navigate to a target indicated by an image in a previously unseen scene. Current methods introduce diverse memory mechanisms which save navigation history to solve…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Hongxin Li , Xu Yang , Yuran Yang , Shuqi Mei , Zhaoxiang Zhang

LLM-powered embodied agents have shown success on conventional object-rearrangement tasks, but providing personalized assistance that leverages user-specific knowledge from past interactions presents new challenges. We investigate these…

Computation and Language · Computer Science 2026-02-16 Taeyoon Kwon , Dongwook Choi , Hyojun Kim , Sunghwan Kim , Seungjun Moon , Beong-woo Kwak , Kuan-Hao Huang , Jinyoung Yeo

Modern robots vary significantly in shape, size, and sensor configurations used to perceive and interact with their environments. However, most navigation policies are embodiment-specific--a policy trained on one robot typically fails to…

We introduce Blueprint-Bench, a benchmark designed to evaluate spatial reasoning capabilities in AI models through the task of converting apartment photographs into accurate 2D floor plans. While the input modality (photographs) is well…

Artificial Intelligence · Computer Science 2025-10-01 Lukas Petersson , Axel Backlund , Axel Wennstöm , Hanna Petersson , Callum Sharrock , Arash Dabiri

Safe visual navigation is critical for indoor mobile robots operating in cluttered environments. Existing benchmarks, however, often neglect collisions or are designed for outdoor scenarios, making them unsuitable for indoor visual…

Robotics · Computer Science 2026-03-05 Jaewon Lee , Jaeseok Heo , Gunmin Lee , Howoong Jun , Jeongwoo Oh , Songhwai Oh

Embodied artificial intelligence emphasizes the role of an agent's body in generating human-like behaviors. The recent efforts on EmbodiedAI pay a lot of attention to building up machine learning models to possess perceiving, planning, and…

Artificial Intelligence · Computer Science 2024-10-15 Chen Gao , Baining Zhao , Weichen Zhang , Jinzhu Mao , Jun Zhang , Zhiheng Zheng , Fanhang Man , Jianjie Fang , Zile Zhou , Jinqiang Cui , Xinlei Chen , Yong Li

Navigation is one of the fundamental tasks for automated exploration in Virtual Reality (VR). Existing technologies primarily focus on path optimization in 360-degree image datasets and 3D simulators, which cannot be directly applied to…

Software Engineering · Computer Science 2026-01-07 Xue Qin , Matthew DiGiovanni

Embodied AI is a crucial frontier in robotics, capable of planning and executing action sequences for robots to accomplish long-horizon tasks in physical environments. In this work, we introduce EmbodiedGPT, an end-to-end multi-modal…

Robotics · Computer Science 2023-09-15 Yao Mu , Qinglong Zhang , Mengkang Hu , Wenhai Wang , Mingyu Ding , Jun Jin , Bin Wang , Jifeng Dai , Yu Qiao , Ping Luo

Safe, agile, and socially compliant multi-robot navigation in cluttered and constrained environments remains a critical challenge. This is especially difficult with self-interested agents with unique, unknown priorities in decentralized…

Robotics · Computer Science 2026-05-12 Vagul Mahadevan , Shangtong Zhang , Rohan Chandra

Commanding a robot to navigate with natural language instructions is a long-term goal for grounded language understanding and robotics. But the dominant language is English, according to previous studies on vision-language navigation (VLN).…

Computation and Language · Computer Science 2020-12-08 An Yan , Xin Eric Wang , Jiangtao Feng , Lei Li , William Yang Wang

Visual grounding in 3D is the key for embodied agents to localize language-referred objects in open-world environments. However, existing benchmarks are limited to indoor focus, single-platform constraints, and small scale. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Rong Li , Yuhao Dong , Tianshuai Hu , Ao Liang , Youquan Liu , Dongyue Lu , Liang Pan , Lingdong Kong , Junwei Liang , Ziwei Liu

Embodied navigation holds significant promise for real-world applications such as last-mile delivery. However, most existing approaches are confined to either indoor or outdoor environments and rely heavily on strong assumptions, such as…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Yuxiang Zhao , Yirong Yang , Yanqing Zhu , Yanfen Shen , Chiyu Wang , Zhining Gu , Pei Shi , Wei Guo , Mu Xu