English
Related papers

Related papers: Extending Embodied Question Answering from Percept…

200 papers

Recent advances in deep thinking models have demonstrated remarkable reasoning capabilities on mathematical and coding tasks. However, their effectiveness in embodied domains which require continuous interaction with environments through…

Computation and Language · Computer Science 2025-05-15 Wenqi Zhang , Mengna Wang , Gangao Liu , Xu Huixin , Yiwei Jiang , Yongliang Shen , Guiyang Hou , Zhe Zheng , Hang Zhang , Xin Li , Weiming Lu , Peng Li , Yueting Zhuang

Embodied Question Answering (EQA) is a relatively new task where an agent is asked to answer questions about its environment from egocentric perception. EQA makes the fundamental assumption that every question, e.g., "what color is the…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Licheng Yu , Xinlei Chen , Georgia Gkioxari , Mohit Bansal , Tamara L. Berg , Dhruv Batra

In this paper, we propose a novel Knowledge-based Embodied Question Answering (K-EQA) task, in which the agent intelligently explores the environment to answer various questions with the knowledge. Different from explicitly specifying the…

Robotics · Computer Science 2021-09-17 Sinan Tan , Mengmeng Ge , Di Guo , Huaping Liu , Fuchun Sun

As embodied models become powerful, humans will collaborate with multiple embodied AI agents at their workplace or home in the future. To ensure better communication between human users and the multi-agent system, it is crucial to interpret…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Kangsan Kim , Yanlai Yang , Suji Kim , Woongyeong Yeo , Youngwan Lee , Mengye Ren , Sung Ju Hwang

3D Scene Question Answering (3D SQA) represents an interdisciplinary task that integrates 3D visual perception and natural language processing, empowering intelligent agents to comprehend and interact with complex 3D environments. Recent…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Zechuan Li , Hongshan Yu , Yihao Ding , Yan Li , Yong He , Naveed Akhtar

Recent advances in embodied AI highlight the potential of vision language models (VLMs) as agents capable of perception, reasoning, and interaction in complex environments. However, top-performing systems rely on large-scale models that are…

Vision Language Models (VLMs) demonstrate significant potential as embodied AI agents for various mobility applications. However, a standardized, closed-loop benchmark for evaluating their spatial reasoning and sequential decision-making…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Weizhen Wang , Chenda Duan , Zhenghao Peng , Yuxin Liu , Bolei Zhou

The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However, current datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Ronghao Dang , Yuqian Yuan , Wenqi Zhang , Yifei Xin , Boqiang Zhang , Long Li , Liuyi Wang , Qinyang Zeng , Xin Li , Lidong Bing

Embodied AI has developed rapidly in recent years, but it is still mainly deployed in laboratories, with various distortions in the Real-world limiting its application. Traditionally, Image Quality Assessment (IQA) methods are applied to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Chunyi Li , Jiaohao Xiao , Jianbo Zhang , Farong Wen , Zicheng Zhang , Yuan Tian , Xiangyang Zhu , Xiaohong Liu , Zhengxue Cheng , Weisi Lin , Guangtao Zhai

We propose a new task to benchmark scene understanding of embodied agents: Situated Question Answering in 3D Scenes (SQA3D). Given a scene context (e.g., 3D scan), SQA3D requires the tested agent to first understand its situation (position,…

Computer Vision and Pattern Recognition · Computer Science 2023-04-14 Xiaojian Ma , Silong Yong , Zilong Zheng , Qing Li , Yitao Liang , Song-Chun Zhu , Siyuan Huang

While video large language models (Video-LLMs) excel in understanding slow-paced, real-world egocentric videos, their capabilities in high-velocity, information-dense virtual environments remain under-explored. Existing benchmarks focus on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Jianzhe Ma , Zhonghao Cao , Shangkui Chen , Yichen Xu , Wenxuan Wang , Qin Jin

Embodied AI aims to develop intelligent systems with physical forms capable of perceiving, decision-making, acting, and learning in real-world environments, providing a promising way to Artificial General Intelligence (AGI). Despite decades…

Robotics · Computer Science 2025-08-15 Wenlong Liang , Rui Zhou , Yang Ma , Bing Zhang , Songlin Li , Yijia Liao , Ping Kuang

Embodied intelligence aims to enable robots to learn, reason, and generalize robustly across complex real-world environments. However, existing approaches often struggle with partial observability, fragmented spatial reasoning, and…

Embodied agents are expected to perform more complicated tasks in an interactive environment, with the progress of Embodied AI in recent years. Existing embodied tasks including Embodied Referring Expression (ERE) and other QA-form tasks…

Robotics · Computer Science 2023-10-18 Qie Sima , Sinan Tan , Huaping Liu

Domain-specific quantitative reasoning remains a major challenge for large language models (LLMs), especially in fields requiring expert knowledge and complex question answering (QA). In this work, we propose Expert Question Decomposition…

Computation and Language · Computer Science 2025-10-03 Mengyu Wang , Sotirios Sabanis , Miguel de Carvalho , Shay B. Cohen , Tiejun Ma

Vision-language navigation requires agents to reason and act under constraints of embodiment. While vision-language models (VLMs) demonstrate strong generalization, current benchmarks provide limited understanding of how embodiment -- i.e.,…

Robotics · Computer Science 2025-12-23 Tin Stribor Sohn , Maximilian Dillitzer , Jason J. Corso , Eric Sax

We present and tackle the problem of Embodied Question Answering (EQA) with Situational Queries (S-EQA) in a household environment. Unlike prior EQA work tackling simple queries that directly reference target objects and properties ("What…

In Embodied Question Answering (EQA), agents must explore and develop a semantic understanding of an unseen environment to answer a situated question with confidence. This problem remains challenging in robotics, due to the difficulties in…

Large language models excel at a wide range of complex tasks. However, enabling general inference in the real world, e.g., for robotics problems, raises the challenge of grounding. We propose embodied language models to directly incorporate…

Embodied AI development significantly lags behind large foundation models due to three critical challenges: (1) lack of systematic understanding of core capabilities needed for Embodied AI, making research lack clear objectives; (2) absence…