English
Related papers

Related papers: Learning Situated Awareness in the Real World

200 papers

Spatial reasoning, the ability to understand spatial relations, causality, and dynamic evolution, is central to human intelligence and essential for real-world applications such as autonomous driving and robotics. Existing studies, however,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Yanguang Zhao , Jie Yang , Shengqiong Wu , Shutong Hu , Hongbo Qiu , Yu Wang , Guijia Zhang , Tan Kai Ze , Hao Fei , Chia-Wen Lin , Mong-Li Lee , Wynne Hsu

Understanding social interaction, which encompasses perceiving numerous and subtle multimodal cues, inferring unobservable mental states and relations, and dynamically predicting others' behavior, is the foundation for achieving…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Fanqi Kong , Weiqin Zu , Xinyu Chen , Yaodong Yang , Song-Chun Zhu , Xue Feng

LMMs have shown impressive visual understanding capabilities, with the potential to be applied in agents, which demand strong reasoning and planning abilities. Nevertheless, existing benchmarks mostly assess their reasoning abilities in…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Miaosen Zhang , Qi Dai , Yifan Yang , Jianmin Bao , Dongdong Chen , Kai Qiu , Chong Luo , Xin Geng , Baining Guo

Applications like personal assistants need to be aware ofthe user's context, e.g., where they are, what they are doing, and with whom. Context information is usually inferred from sensor data, like GPS sensors and accelerometers on the…

Artificial Intelligence · Computer Science 2020-11-20 Qiang Shen , Stefano Teso , Wanyi Zhang , Hao Xu , Fausto Giunchiglia

The rapid evolution of Large Multimodal Models (LMMs) has enabled agents to perform complex digital and physical tasks, yet their deployment as autonomous decision-makers introduces substantial unintentional behavioral safety risks.…

Artificial Intelligence · Computer Science 2026-03-30 Yuxuan Li , Yi Lin , Peng Wang , Shiming Liu , Xuetao Wei

Being able to carry out complicated vision language reasoning tasks in 3D space represents a significant milestone in developing household robots and human-centered embodied AI. In this work, we demonstrate that a critical and distinct…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Yunze Man , Liang-Yan Gui , Yu-Xiong Wang

Spatial intelligence is crucial for vision--language models (VLMs) in the physical world, yet many benchmarks evaluate largely unconstrained scenes where models can exploit 2D shortcuts. We introduce SSI-Bench, a VQA benchmark for spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Chen Yang , Guanxin Lin , Youquan He , Peiyao Chen , Guanghe Liu , Yufan Mo , Zhouyuan Xu , Linhao Wang , Guohui Zhang , Zihang Zhang , Shenxiang Zeng , Chen Wang , Jiansheng Fan

Reliable localization of people is fundamental for service and social robots that must operate in close interaction with humans. State-of-the-art human detectors often rely on RGB-D cameras or costly 3D LiDARs. However, most commercial…

Robotics · Computer Science 2026-04-17 Simone Arreghini , Nicholas Carlotti , Mirko Nava , Antonio Paolillo , Alessandro Giusti

Achieving fine-grained spatio-temporal understanding in videos remains a major challenge for current Video Large Multimodal Models (Video LMMs). Addressing this challenge requires mastering two core capabilities: video referring…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Ye Sun , Hao Zhang , Henghui Ding , Tiehua Zhang , Xingjun Ma , Yu-Gang Jiang

Egocentric video recognition is a natural testbed for diverse interaction reasoning. Due to the large action vocabulary in egocentric video datasets, recent studies usually utilize a two-branch structure for action recognition, ie, one…

Computer Vision and Pattern Recognition · Computer Science 2020-02-11 Xiaohan Wang , Yu Wu , Linchao Zhu , Yi Yang

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jinzhou Tang , Jusheng zhang , Sidi Liu , Waikit Xiu , Qinhan Lv , Xiying Li

Situation awareness is essential for understanding and reasoning about 3D scenes in embodied AI agents. However, existing datasets and benchmarks for situated understanding are limited in data modality, diversity, scale, and task scope. To…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Xiongkun Linghu , Jiangyong Huang , Xuesong Niu , Xiaojian Ma , Baoxiong Jia , Siyuan Huang

Humans build viewpoint-independent cognitive maps through navigation, enabling intuitive reasoning about object permanence and spatial relations. We argue that multimodal large language models (MLLMs), despite extensive video training, lack…

Machine Learning · Computer Science 2025-12-02 Jacob Thompson , Emiliano Garcia-Lopez , Yonatan Bisk

Vision Language Models (VLMs) perform well on standard video tasks but struggle with physics-related reasoning involving motion dynamics and spatial interactions. We present a novel approach to address this gap by translating physical-world…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Xiyang Wu , Zongxia Li , Jihui Jin , Guangyao Shi , Gouthaman KV , Vishnu Raj , Nilotpal Sinha , Jingxi Chen , Fan Du , Dinesh Manocha

Transferring and integrating knowledge across first-person (egocentric) and third-person (exocentric) viewpoints is intrinsic to human intelligence, enabling humans to learn from others and convey insights from their own experiences.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Yuping He , Yifei Huang , Guo Chen , Baoqi Pei , Jilan Xu , Tong Lu , Jiangmiao Pang

We investigate the ability of Vision Language Models (VLMs) to perform visual perspective taking using a new set of visual tasks inspired by established human tests. Our approach leverages carefully controlled scenes in which a single…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Gracjan Góral , Alicja Ziarko , Piotr Miłoś , Michał Nauman , Maciej Wołczyk , Michał Kosiński

Human activity recognition is typically addressed by detecting key concepts like global and local motion, features related to object classes present in the scene, as well as features related to the global context. The next open challenges…

Computer Vision and Pattern Recognition · Computer Science 2018-09-21 Fabien Baradel , Natalia Neverova , Christian Wolf , Julien Mille , Greg Mori

Spatial reasoning is a core aspect of human intelligence that allows perception, inference and planning in 3D environments. However, current vision-language models (VLMs) struggle to maintain geometric coherence and cross-view consistency…

Artificial Intelligence · Computer Science 2025-12-03 Qiyao Xue , Weichen Liu , Shiqi Wang , Haoming Wang , Yuyang Wu , Wei Gao

We investigate the real-time estimation of human situation awareness using observations from a robot teammate with limited visibility. In human factors and human-autonomy teaming, it is recognized that individuals navigate their…

Robotics · Computer Science 2025-02-13 Jack Kolb , Karen M. Feigh

Spatial understanding is fundamental for embodied agents, yet most spatial VLMs and benchmarks remain offline-evaluating post-hoc QA over pre-recorded inputs and overlooking two crucial deployment-critical requirements: long-horizon…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Yuxi Wei , Wei Huang , Qirui Chen , Lu Hou , Xiaojuan Qi