English
Related papers

Related papers: Scaling Spatial Intelligence with Multimodal Found…

200 papers

Multimodal learning, especially large-scale multimodal pre-training, has developed rapidly over the past few years and led to the greatest advances in artificial intelligence (AI). Despite its effectiveness, understanding the underlying…

Neural and Evolutionary Computing · Computer Science 2022-08-18 Haoyu Lu , Qiongyi Zhou , Nanyi Fei , Zhiwu Lu , Mingyu Ding , Jingyuan Wen , Changde Du , Xin Zhao , Hao Sun , Huiguang He , Ji-Rong Wen

While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, visuospatial cognition - reasoning about spatial layouts, relations, and dynamics - remains a significant challenge. Existing models often lack the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Qi Feng

This paper presents preliminary results in the definition of a comprehensive benchmark framework designed to systematically evaluate spatial reasoning capabilities in neural networks, with a particular focus on morphological properties such…

Machine Learning · Computer Science 2025-08-19 Manuela Imbriani , Gina Belmonte , Mieke Massink , Alessandro Tofani , Vincenzo Ciancia

Scaling up model and data size have demonstrated impressive performance improvement over a wide range of tasks. Despite extensive studies on scaling behaviors for general-purpose tasks, medical images exhibit substantial differences from…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Jiarun Liu , Hong-Yu Zhou , Weijian Huang , Hao Yang , Dongning Song , Tao Tan , Yong Liang , Shanshan Wang

Spatial reasoning is a core aspect of human intelligence that allows perception, inference and planning in 3D environments. However, current vision-language models (VLMs) struggle to maintain geometric coherence and cross-view consistency…

Artificial Intelligence · Computer Science 2025-12-03 Qiyao Xue , Weichen Liu , Shiqi Wang , Haoming Wang , Yuyang Wu , Wei Gao

Humans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (LMMs), however, lack of this capability of 3D spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Wufei Ma , Luoxin Ye , Celso M de Melo , Jieneng Chen , Alan Yuille

We explore the scaling behaviors of artificial intelligence to establish practical techniques for training foundation models on high-resolution electro-optical (EO) datasets that exceed the current state-of-the-art scale by orders of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Charith Wickrema , Eliza Mace , Hunter Brown , Heidys Cabrera , Nick Krall , Matthew O'Neill , Shivangi Sarkar , Lowell Weissman , Eric Hughes , Guido Zarrella

Over the past year, the development of large language models (LLMs) has brought spatial intelligence into focus, with much attention on vision-based embodied intelligence. However, spatial intelligence spans a broader range of disciplines…

Despite recent progress in multimodal agentic systems, existing approaches often treat image manipulation and web search as disjoint capabilities, rely heavily on costly reinforcement learning, and lack planning grounded in real…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Yifan Zhang , Liang Hu , Haofeng Sun , Peiyu Wang , Yichen Wei , Shukang Yin , Jiangbo Pei , Wei Shen , Peng Xia , Yi Peng , Tianyidan Xie , Eric Li , Yang Liu , Xuchen Song , Yahui Zhou

Recent advancements in Spatial Intelligence (SI) have predominantly relied on Vision-Language Models (VLMs), yet a critical question remains: does spatial understanding originate from visual encoders or the fundamental reasoning backbone?…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Zhongbin Guo , Zhen Yang , Yushan Li , Xinyue Zhang , Wenyu Gao , Jiacheng Wang , Chengzhi Li , Xiangrui Liu , Ping Jian

Vision-Language Models (VLMs) have recently emerged as powerful tools, excelling in tasks that integrate visual and textual comprehension, such as image captioning, visual question answering, and image-text retrieval. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Ilias Stogiannidis , Steven McDonagh , Sotirios A. Tsaftaris

Large language models (LLMs) and vision-language models (VLMs) have demonstrated remarkable performance across a wide range of tasks and domains. Despite this promise, spatial understanding and reasoning -- a fundamental component of human…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Jiayu Wang , Yifei Ming , Zhenmei Shi , Vibhav Vineet , Xin Wang , Yixuan Li , Neel Joshi

Embodied intelligence, a grand challenge in artificial intelligence, is fundamentally constrained by the limited spatial understanding and reasoning capabilities of current models. Prevailing efforts to address this through enhancing…

Artificial Intelligence · Computer Science 2025-12-19 Zhi Helu , Huang Jingjing , Xu Wang , Xu Yangbin , Zhang Wanyue , Jiang Baoyang , Deng Shirui , Zhu Liang , Li Fangfang , Zhao Tiejun , Lin Yankai , Yao Yuan

Multimodal large language models (MLLMs) have shown remarkable potential in various domains, yet their application in the medical field is hindered by several challenges. General-purpose MLLMs often lack the specialized knowledge required…

Artificial Intelligence · Computer Science 2025-09-29 Guanghao Zhu , Zhitian Hou , Zeyu Liu , Zhijie Sang , Congkai Xie , Hongxia Yang

We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond…

Machine Learning · Computer Science 2026-04-03 Yicheng Zou , Dongsheng Zhu , Lin Zhu , Tong Zhu , Yunhua Zhou , Peiheng Zhou , Xinyu Zhou , Dongzhan Zhou , Zhiwang Zhou , Yuhao Zhou , Bowen Zhou , Zhanping Zhong , Zhijie Zhong , Haiteng Zhao , Penghao Zhao , Xiaomeng Zhao , Zhiyuan Zhao , Yechen Zhang , Jin Zhang , Wenwei Zhang , Hongjie Zhang , Zhuo Zhang , Wenlong Zhang , Bo Zhang , Chao Zhang , Chen Zhang , Yuhang Zang , Fei Yuan , Jiakang Yuan , Jiashuo Yu , Jinhui Yin , Haochen Ye , Qian Yao , Bowen Yang , Danni Yang , Kaichen Yang , Ziang Yan , Jun Xu , Yicheng Xu , Wanghan Xu , Xuenan Xu , Chao Xu , Ruiliang Xu , Shuhao Xing , Long Xing , Xinchen Xie , Ling-I Wu , Zijian Wu , Zhenyu Wu , Lijun Wu , Yue Wu , Jianyu Wu , Wen Wu , Fan Wu , Xilin Wei , Qi Wei , Bingli Wang , Rui Wang , Ziyi Wang , Zun Wang , Yi Wang , Haomin Wang , Yizhou Wang , Lintao Wang , Yiheng Wang , Longjiang Wang , Bin Wang , Jian Tong , Zhongbo Tian , Huanze Tang , Chen Tang , Shixiang Tang , Yu Sun , Qiushi Sun , Xuerui Su , Qisheng Su , Chenlin Su , Demin Song , Jin Shi , Fukai Shang , Yuchen Ren , Pengli Ren , Xiaoye Qu , Yuan Qu , Jiantao Qiu , Yu Qiao , Biqing Qi , Runyu Peng , Tianshuo Peng , Jiahui Peng , Qizhi Pei , Zhuoshi Pan , Linke Ouyang , Wenchang Ning , Yichuan Ma , Zerun Ma , Ningsheng Ma , Runyuan Ma , Chengqi Lyu , Haijun Lv , Han Lv , Lindong Lu , Kuikun Liu , Jiangning Liu , Yuhong Liu , Kai Liu , Hongwei Liu , Zhoumianze Liu , Mengjie Liu , Ziyu Liu , Wenran Liu , Yang Liu , Liwei Liu , Kaiwen Liu , Junyao Lin , Junming Lin , Tianyang Lin , Dahua Lin , Jianze Liang , Linyang Li , Peiji Li , Zonglin Li , Zehao Li , Pengze Li , Guoyan Li , Lingkai Kong , Linglin Jing , Zhenjiang Jin , Feifei Jiang , Qian Jiang , Junhao Huang , Zixian Huang , Haian Huang , Zhouqi Hua , Ermo Hua , Han Hu , Linfeng Hou , Yinan He , Conghui He , Tianyao He , Xu Guo , Qipeng Guo , Aijia Guo , Yuzhe Gu , Lixin Gu , Jingyang Gong , Qiming Ge , Jiaye Ge , Songyang Gao , Jianfei Gao , Xinyu Fang , Caihua fan , Yue Fan , Yanhui Duan , Zichen Ding , Shengyuan Ding , Ning Ding , Xuanlang Dai , Erfei Cui , Ganqu Cui , Pei Chu , Tao Chu , Guangran Cheng , Yu Cheng , Kai Chen , Yongkang Chen , Chiyu Chen , Guanzhou Chen , Qiaosheng Chen , Sitao Chen , Xin Chen , Haojiong Chen , Yicheng Chen , Weihan Cao , Yuhang Cao , Qinglong Cao , Lei Bai

As reasoning models scale rapidly, the essential role of multimodality in human cognition has come into sharp relief, driving a growing need to probe vision-centric cognitive behaviors. Yet, existing multimodal benchmarks either…

Large language models (LLMs) and vision language models (VLMs), such as DeepSeek R1,OpenAI o3, and Gemini 2.5 Pro, have demonstrated remarkable reasoning capabilities across logical inference, problem solving, and decision making. However,…

Artificial Intelligence · Computer Science 2025-11-19 Xiaoxing Lian , Aidong Yang , Jun Zhu , Peng Wang , Yue Zhang

We argue that progress in true multimodal intelligence calls for a shift from reactive, task-driven systems and brute-force long context towards a broader paradigm of supersensing. We frame spatial supersensing as four stages beyond…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Shusheng Yang , Jihan Yang , Pinzhi Huang , Ellis Brown , Zihao Yang , Yue Yu , Shengbang Tong , Zihan Zheng , Yifan Xu , Muhan Wang , Daohan Lu , Rob Fergus , Yann LeCun , Li Fei-Fei , Saining Xie

The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning. A recent line of work explores learning spatial reasoning directly from multi-view images,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Kanghee Lee , Injae Lee , Minseok Kwak , Jungi Hong , Kwonyoung Ryu , Jaesik Park

Multimodal large language models (MLLMs) have shown great potential in perception and interpretation tasks, but their capabilities in predictive reasoning remain under-explored. To address this gap, we introduce a novel benchmark that…

Computer Vision and Pattern Recognition · Computer Science 2023-10-23 Mingwei Zhu , Leigang Sha , Yu Shu , Kangjia Zhao , Tiancheng Zhao , Jianwei Yin