English
Related papers

Related papers: EgoMind: Activating Spatial Cognition through Ling…

200 papers

Video spatial reasoning, which involves inferring the underlying spatial structure from observed video frames, poses a significant challenge for existing Multimodal Large Language Models (MLLMs). This limitation stems primarily from 1) the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Kun Ouyang , Yuanxin Liu , Haoning Wu , Yi Liu , Hao Zhou , Jie Zhou , Fandong Meng , Xu Sun

Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We identify that this limitation stems from a critical gap:…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Hongxing Li , Dingming Li , Zixuan Wang , Yuchen Yan , Hang Wu , Wenqi Zhang , Yongliang Shen , Weiming Lu , Jun Xiao , Yueting Zhuang

Spatial relation reasoning is a crucial task for multimodal large language models (MLLMs) to understand the objective world. However, current benchmarks have issues like relying on bounding boxes, ignoring perspective substitutions, or…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Jingping Liu , Ziyan Liu , Zhedong Cen , Yan Zhou , Yinan Zou , Weiyan Zhang , Haiyun Jiang , Tong Ruan

The rapid advancing of Multimodal Large Language Models (MLLMs) has spurred interest in complex multimodal reasoning tasks in the real-world and virtual environment, which require coordinating multiple abilities, including visual…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Ziyue Wang , Yurui Dong , Fuwen Luo , Minyuan Ruan , Zhili Cheng , Chi Chen , Peng Li , Yang Liu

Large language models (LLMs) have demonstrated strong reasoning capabilities in text-based mathematical problem solving; however, when adapted to visual reasoning tasks, particularly geometric problem solving, their performance…

Artificial Intelligence · Computer Science 2025-10-28 Nannan Shi , Chuanyu Qin , Shipeng Song , Man Luo

Vision-language models (VLMs) have advanced multimodal reasoning but still face challenges in spatial reasoning for 3D scenes and complex object configurations. To address this, we introduce SpatialViLT, an enhanced VLM that integrates…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Chashi Mahiul Islam , Oteo Mamo , Samuel Jacob Chacko , Xiuwen Liu , Weikuan Yu

Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a core challenge. Recent embodied systems often adopt a dual-system paradigm, where System 2 handles high-level reasoning while System 1…

The 180x360 omnidirectional field of view captured by 360-degree cameras enables their use in a wide range of applications such as embodied AI and virtual reality. Although recent advances in multimodal large language models (MLLMs) have…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Zihao Dongfang , Xu Zheng , Ziqiao Weng , Yuanhuiyi Lyu , Danda Pani Paudel , Luc Van Gool , Kailun Yang , Xuming Hu

Multimodal language models (MLMs) perform well on semantic vision-language tasks but fail at spatial reasoning that requires adopting another agent's visual perspective. These errors reflect a persistent egocentric bias and raise questions…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Bridget Leonard , Scott O. Murray

As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Junfei Wu , Jian Guan , Kaituo Feng , Qiang Liu , Shu Wu , Liang Wang , Wei Wu , Tieniu Tan

Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongside spoken words, and respond with empathy that aligns with…

Computation and Language · Computer Science 2026-03-06 Li Zhou , Lutong Yu , You Lyu , Yihang Lin , Zefeng Zhao , Junyi Ao , Yuhao Zhang , Benyou Wang , Haizhou Li

Multimodal large language models (MLLMs) have undergone rapid development in advancing geospatial scene understanding. Recent studies have sought to enhance the reasoning capabilities of remote sensing MLLMs, typically through cold-start…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Di Wang , Shunyu Liu , Wentao Jiang , Fengxiang Wang , Yi Liu , Xiaolei Qin , Zhiming Luo , Chaoyang Zhou , Haonan Guo , Jing Zhang , Bo Du , Dacheng Tao , Liangpei Zhang

Transferring and integrating knowledge across first-person (egocentric) and third-person (exocentric) viewpoints is intrinsic to human intelligence, enabling humans to learn from others and convey insights from their own experiences.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Yuping He , Yifei Huang , Guo Chen , Baoqi Pei , Jilan Xu , Tong Lu , Jiangmiao Pang

Large Language Models (LLMs) have undergone rapid progress, largely attributed to reinforcement learning on complex reasoning tasks. In contrast, while spatial intelligence is fundamental for Vision-Language Models (VLMs) in real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Zijian Song , Xiaoxin Lin , Qiuming Huang , Sihan Qin , Guangrun Wang , Liang Lin

Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoints remains largely unexplored. We study this problem through…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Kunyu Peng , Zhikun Zhou , Kailun Yang , Di Wen , Ruiping Liu , Yufan Chen , Junwei Zheng , Hao Shi , Yi Zhou , M. Saquib Sarfraz , Danda Pani Paudel , Luc Van Gool

We explore leveraging large multi-modal models (LMMs) and text2image models to build a more general embodied agent. LMMs excel in planning long-horizon tasks over symbolic abstractions but struggle with grounding in the physical world,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Zhirui Fang , Ming Yang , Weishuai Zeng , Boyu Li , Junpeng Yue , Ziluo Ding , Xiu Li , Zongqing Lu

Genuine spatial reasoning relies on the capacity to construct and manipulate coherent internal spatial representations, often conceptualized as mental models, rather than merely processing surface linguistic associations. While large…

Artificial Intelligence · Computer Science 2026-03-04 Peiyao Jiang , Zequn Qin , Xi Li

Current state-of-the-art spatial reasoning-enhanced VLMs are trained to excel at spatial visual question answering (VQA). However, we believe that higher-level 3D-aware tasks, such as articulating dynamic scene changes and motion planning,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Chenyang Ma , Kai Lu , Ta-Ying Cheng , Niki Trigoni , Andrew Markham

Multimodal Small-to-Medium sized Language Models (MSLMs) have demonstrated strong capabilities in integrating visual and textual information but still face significant limitations in visual comprehension and mathematical reasoning,…

Machine Learning · Computer Science 2026-01-27 Ashutosh Bajpai , Akshat Bhandari , Akshay Nambi , Tanmoy Chakraborty

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Haoyu Zhang , Meng Liu , Zaijing Li , Haokun Wen , Weili Guan , Yaowei Wang , Liqiang Nie