中文
相关论文

相关论文: Learning Situated Awareness in the Real World

200 篇论文

3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Wufei Ma , Haoyu Chen , Guofeng Zhang , Yu-Cheng Chou , Jieneng Chen , Celso M de Melo , Alan Yuille

Recent advancements in Spatial Intelligence (SI) have predominantly relied on Vision-Language Models (VLMs), yet a critical question remains: does spatial understanding originate from visual encoders or the fundamental reasoning backbone?…

计算机视觉与模式识别 · 计算机科学 2026-01-08 Zhongbin Guo , Zhen Yang , Yushan Li , Xinyue Zhang , Wenyu Gao , Jiacheng Wang , Chengzhi Li , Xiangrui Liu , Ping Jian

As humans move around, performing their daily tasks, they are able to recall where they have positioned objects in their environment, even if these objects are currently out of their sight. In this paper, we aim to mimic this spatial…

计算机视觉与模式识别 · 计算机科学 2025-01-23 Chiara Plizzari , Shubham Goel , Toby Perrett , Jacob Chalk , Angjoo Kanazawa , Dima Damen

As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to…

人工智能 · 计算机科学 2026-05-28 Yunqi Liu , Tong Niu , Zitong Wang , Zhenlong Dai , Yuqi Qing , Weiqiang Wang , Jian Liu

Tracking moving objects is a critical skill for many everyday tasks, such as crossing a busy street, driving a car or catching a ball. Attention is a key cognitive function that supports object tracking; however, our understanding of the…

For human cognitive process, spatial reasoning and perception are closely entangled, yet the nature of this interplay remains underexplored in the evaluation of multimodal large language models (MLLMs). While recent MLLM advancements show…

计算与语言 · 计算机科学 2025-08-28 Chengzu Li , Wenshan Wu , Huanyu Zhang , Qingtao Li , Zeyu Gao , Yan Xia , José Hernández-Orallo , Ivan Vulić , Furu Wei

Existing benchmarks often highlight the remarkable performance achieved by state-of-the-art Multimodal Foundation Models (MFMs) in leveraging temporal context for video understanding. However, how well do the models truly perform visual…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Ziyao Shangguan , Chuhan Li , Yuxuan Ding , Yanan Zheng , Yilun Zhao , Tesca Fitzgerald , Arman Cohan

Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate frame that relates observations across video frames. Yet this signal is largely absent from multimodal LLMs (MLLMs) for video…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Jihan Yang , Zifan Zhao , Xichen Pan , Shusheng Yang , Junyi Zhang , Bingyi Kang , Hu Xu , Saining Xie

Accurately estimating the 3D pose of the camera wearer in egocentric video sequences is crucial to modeling human behavior in virtual and augmented reality applications. The task presents unique challenges due to the limited visibility of…

计算机视觉与模式识别 · 计算机科学 2024-11-08 Luca Scofano , Alessio Sampieri , Edoardo De Matteis , Indro Spinelli , Fabio Galasso

Large Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing "under" or "behind" relationships between only two objects, pose significant…

计算与语言 · 计算机科学 2025-10-14 Shiqi Chen , Tongyao Zhu , Ruochen Zhou , Jinghan Zhang , Siyang Gao , Juan Carlos Niebles , Mor Geva , Junxian He , Jiajun Wu , Manling Li

Recent self-supervised learning (SSL) models trained on human-like egocentric visual inputs substantially underperform on image recognition tasks compared to humans. These models train on raw, uniform visual inputs collected from…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Timothy Schaumlöffel , Arthur Aubret , Gemma Roig , Jochen Triesch

Humans excel at spatial-temporal reasoning, effortlessly interpreting dynamic visual events from an egocentric viewpoint. However, whether multimodal large language models (MLLMs) can similarly understand the 4D world remains uncertain.…

计算机视觉与模式识别 · 计算机科学 2025-04-24 Peiran Wu , Yunze Liu , Miao Liu , Junxiao Shen

The rapid development of Multi-modality Large Language Models (MLLMs) has significantly influenced various aspects of industry and daily life, showcasing impressive capabilities in visual perception and understanding. However, these models…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Yinan Sun , Zicheng Zhang , Haoning Wu , Xiaohong Liu , Weisi Lin , Guangtao Zhai , Xiongkuo Min

In real-world scenes, target objects may reside in regions that are not visible. While humans can often infer the locations of occluded objects from context and commonsense knowledge, this capability remains a major challenge for…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Posheng Chen , Powen Cheng , Gueter Josmy Faure , Hung-Ting Su , Winston H. Hsu

Research indicates that humans can mistakenly assume that robots and humans have the same field of view, possessing an inaccurate mental model of robots. This misperception may lead to failures during human-robot collaboration tasks where…

机器人学 · 计算机科学 2026-03-09 Hong Wang , Ridhima Phatak , James Ocampo , Zhao Han

Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely assess fine-grained reasoning about human speech. Many tasks remain visually solvable or only…

Incorporating the physical environment is essential for a complete understanding of human behavior in unconstrained every-day tasks. This is especially important in ego-centric tasks where obtaining 3 dimensional information is both…

计算机视觉与模式识别 · 计算机科学 2018-07-30 Mickey Li , Noyan Songur , Pavel Orlov , Stefan Leutenegger , A Aldo Faisal

World models aim to understand, remember, and predict dynamic visual environments, yet a unified benchmark for evaluating their fundamental abilities remains lacking. To address this gap, we introduce MIND, the first open-domain closed-loop…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Yixuan Ye , Xuanyu Lu , Yuxin Jiang , Yuchao Gu , Rui Zhao , Qiwei Liang , Jiachun Pan , Fengda Zhang , Weijia Wu , Alex Jinpeng Wang

Emotion recognition plays a crucial role in various domains of human-robot interaction. In long-term interactions with humans, robots need to respond continuously and accurately, however, the mainstream emotion recognition methods mostly…

人机交互 · 计算机科学 2024-01-23 Zihan Lin , Francisco Cruz , Eduardo Benitez Sandoval

This study examined whether a single ceiling-mounted camera could be used to capture fine-grained learning behaviours in co-located practical learning. In undergraduate nursing simulations, teachers first identified seven observable…

人机交互 · 计算机科学 2026-03-17 Xinyu Li , Linxuan Zhao , Roberto Martinez-Maldonado , Dragan Gasevic , Lixiang Yan