English
Related papers

Related papers: Scaling Spatial Intelligence with Multimodal Found…

200 papers

As spatial intelligence becomes an increasingly important capability for foundation models, it remains unclear whether large language models' (LLMs) performance on spatial reasoning benchmarks reflects structured internal spatial…

Computation and Language · Computer Science 2026-03-30 Jiyuan An , Liner Yang , Mengyan Wang , Luming Lu , Weihua An , Erhong Yang

Large vision-language models (VLMs) still struggle with reliable 3D spatial reasoning, a core capability for embodied and physical AI systems. This limitation arises from their inability to capture fine-grained 3D geometry and spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Jian Zhang , Shijie Zhou , Bangya Liu , Achuta Kadambi , Zhiwen Fan

We introduce Cube Bench, a Rubik's-cube benchmark for evaluating spatial and sequential reasoning in multimodal large language models (MLLMs). The benchmark decomposes performance into five skills: (i) reconstructing cube faces from images…

Computation and Language · Computer Science 2025-12-24 Dhruv Anand , Ehsan Shareghi

Video spatial reasoning, which involves inferring the underlying spatial structure from observed video frames, poses a significant challenge for existing Multimodal Large Language Models (MLLMs). This limitation stems primarily from 1) the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Kun Ouyang , Yuanxin Liu , Haoning Wu , Yi Liu , Hao Zhou , Jie Zhou , Fandong Meng , Xu Sun

Robotic Foundation Models (RFMs) hold great promise as generalist, end-to-end systems for robot control. Yet their ability to generalize across new environments, tasks, and embodiments remains limited. We argue that a major bottleneck lies…

With the increasing integration of visual and textual content in Social Networking Services (SNS), evaluating the multimodal capabilities of Large Language Models (LLMs) is crucial for enhancing user experience, content understanding, and…

Computation and Language · Computer Science 2025-12-16 Hongcheng Guo , Zheyong Xie , Shaosheng Cao , Boyang Wang , Weiting Liu , Anjie Le , Lei Li , Zhoujun Li

Spatial reasoning is a key aspect of cognitive psychology and remains a bottleneck for current vision-language models (VLMs). While extensive research has aimed to evaluate or improve VLMs' understanding of basic spatial relations, such as…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Mengdi Jia , Zekun Qi , Shaochen Zhang , Wenyao Zhang , Xinqiang Yu , Jiawei He , He Wang , Li Yi

Recent vision-language (VL) models are powerful, but can they reliably distinguish "right" from "left"? We curate three new corpora to quantify model comprehension of such basic spatial relations. These tests isolate spatial reasoning more…

Computation and Language · Computer Science 2023-10-31 Amita Kamath , Jack Hessel , Kai-Wei Chang

In recent years, a plethora of open-source foundation models have emerged, achieving remarkable progress in some widely attended fields, with performance being quite close to that of closed-source models. However, in high-value but more…

Machine Learning · Computer Science 2025-08-26 Lei Bai , Zhongrui Cai , Yuhang Cao , Maosong Cao , Weihan Cao , Chiyu Chen , Haojiong Chen , Kai Chen , Pengcheng Chen , Ying Chen , Yongkang Chen , Yu Cheng , Pei Chu , Tao Chu , Erfei Cui , Ganqu Cui , Long Cui , Ziyun Cui , Nianchen Deng , Ning Ding , Nanqing Dong , Peijie Dong , Shihan Dou , Sinan Du , Haodong Duan , Caihua Fan , Ben Gao , Changjiang Gao , Jianfei Gao , Songyang Gao , Yang Gao , Zhangwei Gao , Jiaye Ge , Qiming Ge , Lixin Gu , Yuzhe Gu , Aijia Guo , Qipeng Guo , Xu Guo , Conghui He , Junjun He , Yili Hong , Siyuan Hou , Caiyu Hu , Hanglei Hu , Jucheng Hu , Ming Hu , Zhouqi Hua , Haian Huang , Junhao Huang , Xu Huang , Zixian Huang , Zhe Jiang , Lingkai Kong , Linyang Li , Peiji Li , Pengze Li , Shuaibin Li , Tianbin Li , Wei Li , Yuqiang Li , Dahua Lin , Junyao Lin , Tianyi Lin , Zhishan Lin , Hongwei Liu , Jiangning Liu , Jiyao Liu , Junnan Liu , Kai Liu , Kaiwen Liu , Kuikun Liu , Shichun Liu , Shudong Liu , Wei Liu , Xinyao Liu , Yuhong Liu , Zhan Liu , Yinquan Lu , Haijun Lv , Hongxia Lv , Huijie Lv , Qitan Lv , Ying Lv , Chengqi Lyu , Chenglong Ma , Jianpeng Ma , Ren Ma , Runmin Ma , Runyuan Ma , Xinzhu Ma , Yichuan Ma , Zihan Ma , Sixuan Mi , Junzhi Ning , Wenchang Ning , Xinle Pang , Jiahui Peng , Runyu Peng , Yu Qiao , Jiantao Qiu , Xiaoye Qu , Yuan Qu , Yuchen Ren , Fukai Shang , Wenqi Shao , Junhao Shen , Shuaike Shen , Chunfeng Song , Demin Song , Diping Song , Chenlin Su , Weijie Su , Weigao Sun , Yu Sun , Qian Tan , Cheng Tang , Huanze Tang , Kexian Tang , Shixiang Tang , Jian Tong , Aoran Wang , Bin Wang , Dong Wang , Lintao Wang , Rui Wang , Weiyun Wang , Wenhai Wang , Jiaqi Wang , Yi Wang , Ziyi Wang , Ling-I Wu , Wen Wu , Yue Wu , Zijian Wu , Linchen Xiao , Shuhao Xing , Chao Xu , Huihui Xu , Jun Xu , Ruiliang Xu , Wanghan Xu , GanLin Yang , Yuming Yang , Haochen Ye , Jin Ye , Shenglong Ye , Jia Yu , Jiashuo Yu , Jing Yu , Fei Yuan , Yuhang Zang , Bo Zhang , Chao Zhang , Chen Zhang , Hongjie Zhang , Jin Zhang , Qiaosheng Zhang , Qiuyinzhe Zhang , Songyang Zhang , Taolin Zhang , Wenlong Zhang , Wenwei Zhang , Yechen Zhang , Ziyang Zhang , Haiteng Zhao , Qian Zhao , Xiangyu Zhao , Xiangyu Zhao , Bowen Zhou , Dongzhan Zhou , Peiheng Zhou , Yuhao Zhou , Yunhua Zhou , Dongsheng Zhu , Lin Zhu , Yicheng Zou

Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision. Whether similar principles apply to modeling brain activity remains unclear. Here we leveraged a dataset of 3.1 million neurons…

Whole-slide image (WSI) analysis is challenging due to the gigapixel scale of slides and their inherent hierarchical multi-resolution structure. Existing multiple instance learning (MIL) approaches often model WSIs as unordered collections…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Dongqing Xie , Yonghuang Wu

Learning transferable multimodal embeddings for urban environments is challenging because urban understanding is inherently spatial, yet existing datasets and benchmarks lack explicit alignment between street-view images and urban…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Jie Zhang , Xingtong Yu , Yuan Fang , Rudi Stouffs , Zdravko Trivic

Recent advancements in Multimodal Large Language Models (MLLMs), particularly through Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced their reasoning abilities. However, a critical gap persists: these…

Artificial Intelligence · Computer Science 2025-07-14 Inclusion AI , : , Fudong Wang , Jiajia Liu , Jingdong Chen , Jun Zhou , Kaixiang Ji , Lixiang Ru , Qingpei Guo , Ruobing Zheng , Tianqi Li , Yi Yuan , Yifan Mao , Yuting Xiao , Ziping Ma

Brain foundation models have achieved remarkable advances across a wide range of neuroscience tasks. However, most existing models are limited to a single functional modality, restricting their ability to exploit complementary…

Machine Learning · Computer Science 2026-05-18 Hanning Guo , Hanwen Bi , Farah Abdellatif , Andrei Galbenus , Jon. N. Shah , Abigail Morrison , Jürgen Dammers

Multimodal semantic segmentation is a pivotal component of computer vision and typically surpasses unimodal methods by utilizing rich information set from various sources.Current models frequently adopt modality-specific frameworks that…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Bingyu Li , Da Zhang , Zhiyuan Zhao , Junyu Gao , Xuelong Li

Foundation models constitute a significant advancement in computer vision: after a single, albeit costly, training phase, they can address a wide array of tasks. In the field of Earth observation, over 75 remote sensing vision foundation…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Pierre Adorni , Minh-Tan Pham , Stéphane May , Sébastien Lefèvre

Multimodal large language models (MLLMs) have made remarkable progress in either temporal or spatial localization. However, they struggle to perform spatio-temporal video grounding. This limitation stems from two major challenges. Firstly,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Jiankang Wang , Zhihan Zhang , Zhihang Liu , Yang Li , Jiannan Ge , Hongtao Xie , Yongdong Zhang

Large multimodal models extend the impressive capabilities of large language models by integrating multimodal understanding abilities. However, it is not clear how they can emulate the general intelligence and reasoning ability of humans.…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Yew Ken Chia , Vernon Toh Yan Han , Deepanway Ghosal , Lidong Bing , Soujanya Poria

In recent years, researchers have increasingly been interested in how to enable Multimodal Large Language Models (MLLM) to possess spatial understanding and reasoning capabilities. However, most existing methods overlook the importance of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Zixian Liu , Zhaoxi Chen , Liang Pan , Ziwei Liu

Situation awareness is essential for understanding and reasoning about 3D scenes in embodied AI agents. However, existing datasets and benchmarks for situated understanding are limited in data modality, diversity, scale, and task scope. To…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Xiongkun Linghu , Jiangyong Huang , Xuesong Niu , Xiaojian Ma , Baoxiong Jia , Siyuan Huang
‹ Prev 1 8 9 10 Next ›