English
Related papers

Related papers: BridgeEQA: Virtual Embodied Agents for Real Bridge…

200 papers

As the world of agentic artificial intelligence applied to robotics evolves, the need for agents capable of building and retrieving memories and observations efficiently is increasing. Robots operating in complex environments must build…

Robotics · Computer Science 2026-04-21 Paolo Riva , Leonardo Gargani , Matteo Frosi , Matteo Matteucci

Embodied Question Answering (EQA) serves as a benchmark task to evaluate the capability of robots to navigate within novel environments and identify objects in response to human queries. However, existing EQA methods often rely on simulated…

Robotics · Computer Science 2024-10-15 Koya Sakamoto , Daichi Azuma , Taiki Miyanishi , Shuhei Kurita , Motoaki Kawanabe

We consider the problem of Embodied Question Answering (EQA), which refers to settings where an embodied agent such as a robot needs to actively explore an environment to gather information until it is confident about the answer to a…

Robotics · Computer Science 2024-07-09 Allen Z. Ren , Jaden Clark , Anushri Dixit , Masha Itkina , Anirudha Majumdar , Dorsa Sadigh

In this paper, we propose a novel Knowledge-based Embodied Question Answering (K-EQA) task, in which the agent intelligently explores the environment to answer various questions with the knowledge. Different from explicitly specifying the…

Robotics · Computer Science 2021-09-17 Sinan Tan , Mengmeng Ge , Di Guo , Huaping Liu , Fuchun Sun

Embodied Question Answering (EQA) requires an agent to interpret language, perceive its environment, and navigate within 3D scenes to produce responses. Existing EQA benchmarks assume that every question must be answered, but embodied…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Tao Wu , Chuhao Zhou , Guangyu Zhao , Haozhi Cao , Yewen Pu , Jianfei Yang

Interactive and embodied tasks pose at least two fundamental challenges to existing Vision & Language (VL) models, including 1) grounding language in trajectories of actions and observations, and 2) referential disambiguation. To tackle…

We present and tackle the problem of Embodied Question Answering (EQA) with Situational Queries (S-EQA) in a household environment. Unlike prior EQA work tackling simple queries that directly reference target objects and properties ("What…

Embodied artificial intelligence emphasizes the role of an agent's body in generating human-like behaviors. The recent efforts on EmbodiedAI pay a lot of attention to building up machine learning models to possess perceiving, planning, and…

Artificial Intelligence · Computer Science 2024-10-15 Chen Gao , Baining Zhao , Weichen Zhang , Jinzhu Mao , Jun Zhang , Zhiheng Zheng , Fanhang Man , Jianjie Fang , Zile Zhou , Jinqiang Cui , Xinlei Chen , Yong Li

Visual understanding requires interpreting both natural scenes and the textual information that appears within them, motivating tasks such as Visual Question Answering (VQA). However, current VQA benchmarks overlook scenarios with visually…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Jianing An , Luyang Jiang , Jie Luo , Wenjun Wu , Lei Huang

Structured scene representations are a core component of embodied agents, helping to consolidate raw sensory streams into readable, modular, and searchable formats. Due to their high computational overhead, many approaches build such…

Artificial Intelligence · Computer Science 2025-06-03 Muhammad Qasim Ali , Saeejith Nair , Alexander Wong , Yuchen Cui , Yuhao Chen

Image Quality Assessment (IQA) of User-Generated Content (UGC) is a critical technique for human Quality of Experience (QoE). However, does the the image quality of Robot-Generated Content (RGC) demonstrate traits consistent with the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Jianbo Zhang , Chunyi Li , Jie Hao , Jun Jia , Huiyu Duan , Guoquan Zheng , Liang Yuan , Guangtao Zhai

With the surge in the development of large language models, embodied intelligence has attracted increasing attention. Nevertheless, prior works on embodied intelligence typically encode scene or historical memory in an unimodal manner,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Yang Liu , Xinshuai Song , Kaixuan Jiang , Weixing Chen , Jingzhou Luo , Guanbin Li , Liang Lin

An embodied task such as embodied question answering (EmbodiedQA), requires an agent to explore the environment and collect clues to answer a given question that related with specific objects in the scene. The solution of such task usually…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Yang Wu , Shirui Feng , Guanbin Li , Liang Lin

Large vision-language models have recently demonstrated impressive performance in planning and control tasks, driving interest in their application to real-world robotics. However, deploying these models for reasoning in embodied contexts…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Karmesh Yadav , Yusuf Ali , Gunshi Gupta , Yarin Gal , Zsolt Kira

Multimodal vision-language models (VLMs) continue to achieve ever-improving scores on chart understanding benchmarks. Yet, we find that this progress does not fully capture the breadth of visual reasoning capabilities essential for…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Kushin Mukherjee , Donghao Ren , Dominik Moritz , Yannick Assogba

Recent advances in embodied AI highlight the potential of vision language models (VLMs) as agents capable of perception, reasoning, and interaction in complex environments. However, top-performing systems rely on large-scale models that are…

Egocentric augmented reality devices such as wearable glasses passively capture visual data as a human wearer tours a home environment. We envision a scenario wherein the human communicates with an AI agent powering such a device by asking…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Samyak Datta , Sameer Dharur , Vincent Cartillier , Ruta Desai , Mukul Khanna , Dhruv Batra , Devi Parikh

Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. We introduce FlowVQA, a novel benchmark aimed at assessing the capabilities of visual…

Computation and Language · Computer Science 2024-07-01 Shubhankar Singh , Purvi Chaurasia , Yerram Varun , Pranshu Pandya , Vatsal Gupta , Vivek Gupta , Dan Roth

We propose a new task to benchmark scene understanding of embodied agents: Situated Question Answering in 3D Scenes (SQA3D). Given a scene context (e.g., 3D scan), SQA3D requires the tested agent to first understand its situation (position,…

Computer Vision and Pattern Recognition · Computer Science 2023-04-14 Xiaojian Ma , Silong Yong , Zilong Zheng , Qing Li , Yitao Liang , Song-Chun Zhu , Siyuan Huang

We introduce WearVQA, the first benchmark specifically designed to evaluate the Visual Question Answering (VQA) capabilities of multi-model AI assistant on wearable devices like smart glasses. Unlike prior benchmarks that focus on…