中文
相关论文

相关论文: Scene Graph for Embodied Exploration in Cluttered …

200 篇论文

Embodied AI is one of the most popular studies in artificial intelligence and robotics, which can effectively improve the intelligence of real-world agents (i.e. robots) serving human beings. Scene knowledge is important for an agent to…

人工智能 · 计算机科学 2024-05-14 Song Yaoxian , Sun Penglei , Liu Haoyu , Li Zhixu , Song Wei , Xiao Yanghua , Zhou Xiaofang

3D multimodal question answering (MQA) plays a crucial role in scene understanding by enabling intelligent agents to comprehend their surroundings in 3D environments. While existing research has primarily focused on indoor household tasks…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Penglei Sun , Yaoxian Song , Xiang Liu , Xiaofei Yang , Qiang Wang , Tiefeng Li , Yang Yang , Xiaowen Chu

Careful robot manipulation in every-day cluttered environments requires an accurate understanding of the 3D scene, in order to grasp and place objects stably and reliably and to avoid colliding with other objects. In general, we must…

机器人学 · 计算机科学 2025-11-11 Aditya Agarwal , Gaurav Singh , Bipasha Sen , Tomás Lozano-Pérez , Leslie Pack Kaelbling

We consider the problem of Embodied Question Answering (EQA), which refers to settings where an embodied agent such as a robot needs to actively explore an environment to gather information until it is confident about the answer to a…

机器人学 · 计算机科学 2024-07-09 Allen Z. Ren , Jaden Clark , Anushri Dixit , Masha Itkina , Anirudha Majumdar , Dorsa Sadigh

An embodied task such as embodied question answering (EmbodiedQA), requires an agent to explore the environment and collect clues to answer a given question that related with specific objects in the scene. The solution of such task usually…

计算机视觉与模式识别 · 计算机科学 2021-10-19 Yang Wu , Shirui Feng , Guanbin Li , Liang Lin

Embodied Question Answering (EQA) combines visual scene understanding, goal-directed exploration, spatial and temporal reasoning under partial observability. A central challenge is to confine physical search to question-relevant subspaces…

机器人学 · 计算机科学 2026-02-18 Haochen Zhang , Nirav Savaliya , Faizan Siddiqui , Enna Sachdeva

Large language models (LLMs) have grown in popularity due to their natural language interface and pre trained knowledge, leading to rapidly increasing success in question-answering (QA) tasks. More recently, multi-agent systems with…

机器学习 · 计算机科学 2024-10-21 Bhrij Patel , Vishnu Sashank Dorbala , Amrit Singh Bedi , Dinesh Manocha

Active sensing and planning in unknown, cluttered environments is an open challenge for robots intending to provide home service, search and rescue, narrow-passage inspection, and medical assistance. Although many active sensing methods…

机器人学 · 计算机科学 2022-08-25 Hanwen Ren , Ahmed H. Qureshi

Recent advances in Large Language Models (LLMs) have helped facilitate exciting progress for robotic planning in real, open-world environments. 3D scene graphs (3DSGs) offer a promising environment representation for grounding such…

机器人学 · 计算机科学 2024-11-01 Meghan Booker , Grayson Byrd , Bethany Kemp , Aurora Schmidt , Corban Rivera

Embodied Question Answering (EQA) is a recently proposed task, where an agent is placed in a rich 3D environment and must act based solely on its egocentric input to answer a given question. The desired outcome is that the agent learns to…

计算机视觉与模式识别 · 计算机科学 2019-08-15 Cătălina Cangea , Eugene Belilovsky , Pietro Liò , Aaron Courville

We are witnessing significant progress on perception models, specifically those trained on large-scale internet images. However, efficiently generalizing these perception models to unseen embodied tasks is insufficiently studied, which will…

机器人学 · 计算机科学 2023-03-21 Ya Jing , Tao Kong

We introduce the novel task of interactive scene exploration, wherein robots autonomously explore environments and produce an action-conditioned scene graph (ACSG) that captures the structure of the underlying environment. The ACSG accounts…

机器人学 · 计算机科学 2024-10-10 Hanxiao Jiang , Binghao Huang , Ruihai Wu , Zhuoran Li , Shubham Garg , Hooshang Nayyeri , Shenlong Wang , Yunzhu Li

The rise of embodied AI applications has enabled robots to perform complex tasks which require a sophisticated understanding of their environment. To enable successful robot operation in such settings, maps must be constructed so that they…

机器人学 · 计算机科学 2025-04-07 Cody Simons , Aritra Samanta , Amit K. Roy-Chowdhury , Konstantinos Karydis

Visual question answering is concerned with answering free-form questions about an image. Since it requires a deep linguistic understanding of the question and the ability to associate it with various objects that are present in the image,…

机器学习 · 计算机科学 2020-07-03 Marcel Hildebrandt , Hang Li , Rajat Koner , Volker Tresp , Stephan Günnemann

In this work, we propose an evaluation protocol for examining the performance of robotic manipulation policies in cluttered scenes. Contrary to prior works, we approach evaluation from a psychophysical perspective, therefore we use a…

机器人学 · 计算机科学 2025-12-01 Amir Rasouli , Montgomery Alban , Sajjad Pakdamansavoji , Zhiyuan Li , Zhanguang Zhang , Aaron Wu , Xuan Zhao

Automatic Chart Question Answering (ChartQA) is challenging due to the complex distribution of chart elements with patterns of the underlying data not explicitly displayed in charts. To address this challenge, we design a joint multimodal…

计算与语言 · 计算机科学 2024-08-12 Yue Dai , Soyeon Caren Han , Wei Liu

To fully leverage the capabilities of mobile manipulation robots, it is imperative that they are able to autonomously execute long-horizon tasks in large unexplored environments. While large language models (LLMs) have shown emergent…

机器人学 · 计算机科学 2024-08-26 Daniel Honerkamp , Martin Büchner , Fabien Despinoy , Tim Welschehold , Abhinav Valada

We introduce ClutterGen, a physically compliant simulation scene generator capable of producing highly diverse, cluttered, and stable scenes for robot learning. Generating such scenes is challenging as each object must adhere to physical…

机器人学 · 计算机科学 2024-10-08 Yinsen Jia , Boyuan Chen

While current visual captioning models have achieved impressive performance, they often assume that the image is well-captured and provides a complete view of the scene. In real-world scenarios, however, a single image may not offer a good…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Anwen Hu , Shizhe Chen , Liang Zhang , Qin Jin

Dexterous grasping in cluttered scenes presents significant challenges due to diverse object geometries, occlusions, and potential collisions. Existing methods primarily focus on single-object grasping or grasp-pose prediction without…

机器人学 · 计算机科学 2025-09-05 Zeyuan Chen , Qiyang Yan , Yuanpei Chen , Tianhao Wu , Jiyao Zhang , Zihan Ding , Jinzhou Li , Yaodong Yang , Hao Dong