English
Related papers

Related papers: HIS-GPT: Towards 3D Human-In-Scene Multimodal Unde…

200 papers

3D multimodal question answering (MQA) plays a crucial role in scene understanding by enabling intelligent agents to comprehend their surroundings in 3D environments. While existing research has primarily focused on indoor household tasks…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Penglei Sun , Yaoxian Song , Xiang Liu , Xiaofei Yang , Qiang Wang , Tiefeng Li , Yang Yang , Xiaowen Chu

We propose a new task to benchmark scene understanding of embodied agents: Situated Question Answering in 3D Scenes (SQA3D). Given a scene context (e.g., 3D scan), SQA3D requires the tested agent to first understand its situation (position,…

Computer Vision and Pattern Recognition · Computer Science 2023-04-14 Xiaojian Ma , Silong Yong , Zilong Zheng , Qing Li , Yitao Liang , Song-Chun Zhu , Siyuan Huang

Learning generic skills for humanoid robots interacting with 3D scenes by mimicking human data is a key research challenge with significant implications for robotics and real-world applications. However, existing methodologies and…

Robotics · Computer Science 2024-12-24 Yun Liu , Bowen Yang , Licheng Zhong , He Wang , Li Yi

3D Scene Question Answering (3D SQA) represents an interdisciplinary task that integrates 3D visual perception and natural language processing, empowering intelligent agents to comprehend and interact with complex 3D environments. Recent…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Zechuan Li , Hongshan Yu , Yihao Ding , Yan Li , Yong He , Naveed Akhtar

Situation awareness is essential for understanding and reasoning about 3D scenes in embodied AI agents. However, existing datasets and benchmarks for situated understanding are limited in data modality, diversity, scale, and task scope. To…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Xiongkun Linghu , Jiangyong Huang , Xuesong Niu , Xiaojian Ma , Baoxiong Jia , Siyuan Huang

Learning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characteristics of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality…

Computer Vision and Pattern Recognition · Computer Science 2022-10-19 Zan Wang , Yixin Chen , Tengyu Liu , Yixin Zhu , Wei Liang , Siyuan Huang

Taking over arbitrary tasks like humans do with a mobile service robot in open-world settings requires a holistic scene perception for decision-making and high-level control. This paper presents a human-inspired scene perception model to…

Robotics · Computer Science 2024-07-09 Florenz Graf , Jochen Lindermayr , Birgit Graf , Werner Kraus , Marco F. Huber

Modeling human-scene interactions (HSI) is essential for understanding and simulating everyday human behaviors. Recent approaches utilizing generative modeling have made progress in this domain; however, they are limited in controllability…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Inwoo Hwang , Bing Zhou , Young Min Kim , Jian Wang , Chuan Guo

The rise of vision-language foundation models marks an advancement in bridging the gap between human and machine capabilities in 3D scene reasoning. Existing 3D reasoning benchmarks assume real-time scene accessibility, which is impractical…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Ye Mao , Weixun Luo , Junpeng Jing , Anlan Qiu , Krystian Mikolajczyk

Training embodied agents to understand 3D scenes as humans do requires large-scale data of people meaningfully interacting with diverse environments, yet such data is scarce. Real-world capture is costly and limited to controlled settings,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Nikita Kister , Pradyumna YM , István Sárándi , Jiayi Wang , Anna Khoreva , Gerard Pons-Moll

We propose a new 3D spatial understanding task of 3D Question Answering (3D-QA). In the 3D-QA task, models receive visual information from the entire 3D scene of the rich RGB-D indoor scan and answer the given textual questions about the 3D…

Computer Vision and Pattern Recognition · Computer Science 2022-05-10 Daichi Azuma , Taiki Miyanishi , Shuhei Kurita , Motoaki Kawanabe

Human vision is capable of transforming two-dimensional observations into an egocentric three-dimensional scene understanding, which underpins the ability to translate complex scenes and exhibit adaptive behaviors. This capability, however,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Pei Liu , Hongliang Lu , Haichao Liu , Haipeng Liu , Xin Liu , Ruoyu Yao , Shengbo Eben Li , Jun Ma

Faces and humans are crucial elements in social interaction and are widely included in everyday photos and videos. Therefore, a deep understanding of faces and humans will enable multi-modal assistants to achieve improved response quality…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Lixiong Qin , Shilong Ou , Miaoxuan Zhang , Jiangning Wei , Yuhang Zhang , Xiaoshuai Song , Yuchen Liu , Mei Wang , Weiran Xu

Enabling agents to understand and interact with complex 3D scenes is a fundamental challenge for embodied artificial intelligence systems. While Multimodal Large Language Models (MLLMs) have achieved significant progress in 2D image…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Haoyuan Li , Rui Liu , Hehe Fan , Yi Yang

Human-robot interaction is increasingly moving toward multi-robot, socially grounded environments. Existing systems struggle to integrate multimodal perception, embodied expression, and coordinated decision-making in a unified framework.…

Robotics · Computer Science 2026-03-25 Shaid Hasan , Breenice Lee , Sujan Sarker , Tariq Iqbal

When humans and robotic agents coexist in an environment, scene understanding becomes crucial for the agents to carry out various downstream tasks like navigation and planning. Hence, an agent must be capable of localizing and identifying…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Mrunmai Vivek Phatak , Julian Lorenz , Nico Hörmann , Jörg Hähner , Rainer Lienhart

Real-world human-built environments are highly dynamic, involving multiple humans and their complex interactions with surrounding objects. While 3D geometry modeling of such scenes is crucial for applications like AR/VR, gaming, and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Sandika Biswas , Qianyi Wu , Biplab Banerjee , Hamid Rezatofighi

Humans are in constant contact with the world as they move through it and interact with it. This contact is a vital source of information for understanding 3D humans, 3D scenes, and the interactions between them. In fact, we demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Hongwei Yi , Chun-Hao P. Huang , Dimitrios Tzionas , Muhammed Kocabas , Mohamed Hassan , Siyu Tang , Justus Thies , Michael J. Black

We present a novel framework for animating humans in 3D scenes using 3D Gaussian Splatting (3DGS), a neural scene representation that has recently achieved state-of-the-art photorealistic results for novel-view synthesis but remains…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Aymen Mir , Jian Wang , Riza Alp Guler , Chuan Guo , Gerard Pons-Moll , Bing Zhou

Human-Scene Interaction (HSI) is a vital component of fields like embodied AI and virtual reality. Despite advancements in motion quality and physical plausibility, two pivotal factors, versatile interaction control and the development of a…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Zeqi Xiao , Tai Wang , Jingbo Wang , Jinkun Cao , Wenwei Zhang , Bo Dai , Dahua Lin , Jiangmiao Pang
‹ Prev 1 2 3 10 Next ›