English
Related papers

Related papers: Embodied Navigation at the Art Gallery

200 papers

Training embodied AI agents depends critically on the visual fidelity of simulation environments and the ability to model dynamic humans. Current simulators rely on mesh-based rasterization with limited visual realism, and their support for…

Real-world data collection for embodied agents remains costly and unsafe, calling for scalable, realistic, and simulator-ready 3D environments. However, existing scene-generation systems often rely on rule-based or task-specific pipelines,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Hongchi Xia , Xuan Li , Zhaoshuo Li , Qianli Ma , Jiashu Xu , Ming-Yu Liu , Yin Cui , Tsung-Yi Lin , Wei-Chiu Ma , Shenlong Wang , Shuran Song , Fangyin Wei

Image-goal navigation aims to steer an agent towards the goal location specified by an image. Most prior methods tackle this task by learning a navigation policy, which extracts visual features of goal and observation images, compares their…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Pengna Li , Kangyi Wu , Jingwen Fu , Sanping Zhou

Embodied AI and robotic systems increasingly depend on scalable, diverse, and physically grounded 3D content for simulation-based training and real-world deployment. While 3D generative modeling has advanced rapidly, embodied applications…

Robotics · Computer Science 2026-05-11 Tianwei Ye , Yifan Mao , Minwen Liao , Jian Liu , Chunchao Guo , Dazhao Du , Quanxin Shou , Fangqi Zhu , Song Guo

Learning to navigate in complex environments with dynamic elements is an important milestone in developing AI agents. In this work we formulate the navigation question as a reinforcement learning problem and show that data efficiency and…

Constructing a physically realistic and accurately scaled simulated 3D world is crucial for the training and evaluation of embodied intelligence tasks. The diversity, realism, low cost accessibility and affordability of 3D data assets are…

Robotics · Computer Science 2025-06-17 Xinjie Wang , Liu Liu , Yu Cao , Ruiqi Wu , Wenkang Qin , Dehui Wang , Wei Sui , Zhizhong Su

Network visualizations are commonly used to analyze relationships in various contexts. To efficiently explore a network visualization, the user needs to quickly navigate to different parts of the network and analyze local details. Recent…

Human-Computer Interaction · Computer Science 2023-03-29 Helen H. Huang , Hanspeter Pfister , Yalong Yang

Progress in Embodied AI has made it possible for end-to-end-trained agents to navigate in photo-realistic environments with high-level reasoning and zero-shot or language-conditioned behavior, but benchmarks are still dominated by…

Embodied agents for creative tasks like photography must bridge the semantic gap between high-level language commands and geometric control. We introduce PhotoAgent, an agent that achieves this by integrating Large Multimodal Models (LMMs)…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Lirong Che , Zhenfeng Gan , Yanbo Chen , Junbo Tan , Xueqian Wang

We present a semantically rich graph representation for indoor robotic navigation. Our graph representation encodes: semantic locations such as offices or corridors as nodes, and navigational behaviors such as enter office or cross a…

Artificial Intelligence · Computer Science 2018-03-13 Gabriel Sepulveda , Juan Carlos Niebles , Alvaro Soto

Communication between embodied AI agents has received increasing attention in recent years. Despite its use, it is still unclear whether the learned communication is interpretable and grounded in perception. To study the grounding of…

Computer Vision and Pattern Recognition · Computer Science 2021-10-13 Shivansh Patel , Saim Wani , Unnat Jain , Alexander Schwing , Svetlana Lazebnik , Manolis Savva , Angel X. Chang

Embodied agents are expected to assist humans by actively exploring unknown environments and reasoning about spatial contexts. When deployed in real life, agents often face sequential tasks where each new task follows the completion of the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Zhongyi Cai , Yi Du , Chen Wang , Yu Kong

A Scene, represented visually using different formats such as RGB-D, LiDAR scan, keypoints, rectangular, spherical, multi-views, etc., contains information implicitly embedded relevant to applications such as scene indexing, vision-based…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Preeti Meena , Himanshu Kumar , Sandeep Yadav

Developing visual perception models for active agents and sensorimotor control are cumbersome to be done in the physical world, as existing algorithms are too slow to efficiently learn in real-time and robots are fragile and costly. This…

Artificial Intelligence · Computer Science 2018-09-03 Fei Xia , Amir Zamir , Zhi-Yang He , Alexander Sax , Jitendra Malik , Silvio Savarese

In this work, we study the problem of Embodied Referring Expression Grounding, where an agent needs to navigate in a previously unseen environment and localize a remote object described by a concise high-level natural language instruction.…

Computer Vision and Pattern Recognition · Computer Science 2022-12-05 Mingxiao Li , Zehao Wang , Tinne Tuytelaars , Marie-Francine Moens

We propose a robotic learning system for autonomous exploration and navigation in unexplored environments. We are motivated by the idea that even an unseen environment may be familiar from previous experiences in similar environments. The…

Robotics · Computer Science 2022-11-24 Huangying Zhan , Hamid Rezatofighi , Ian Reid

Developing embodied agents in simulation has been a key research topic in recent years. Exciting new tasks, algorithms, and benchmarks have been developed in various simulators. However, most of them assume deaf agents in silent…

Robotics · Computer Science 2023-09-19 Ruohan Gao , Hao Li , Gokul Dharan , Zhuzhu Wang , Chengshu Li , Fei Xia , Silvio Savarese , Li Fei-Fei , Jiajun Wu

Understanding the geometric relationships between objects in a scene is a core capability in enabling both humans and autonomous agents to navigate in new environments. A sparse, unified representation of the scene topology will allow…

Computer Vision and Pattern Recognition · Computer Science 2022-05-18 Zachary Seymour , Niluthpol Chowdhury Mithun , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

To perform tasks specified by natural language instructions, autonomous agents need to extract semantically meaningful representations of language and map it to visual elements and actions in the environment. This problem is called…

We introduce a learning-based approach for room navigation using semantic maps. Our proposed architecture learns to predict top-down belief maps of regions that lie beyond the agent's field of view while modeling architectural and stylistic…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Medhini Narasimhan , Erik Wijmans , Xinlei Chen , Trevor Darrell , Dhruv Batra , Devi Parikh , Amanpreet Singh
‹ Prev 1 4 5 6 7 8 10 Next ›