中文
相关论文

相关论文: World2Mind: Cognition Toolkit for Allocentric Spat…

200 篇论文

Large Language Models (LLMs) have demonstrated impressive reasoning capabilities, yet they struggle with spatial tasks that require mental simulation, such as mental rotation. This paper investigates whether equipping an LLM with an…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Sergio Y. Hayashi , Nina S. T. Hirata

Structured tables are essential for conveying high-density information in professional domains such as finance, healthcare, and scientific research. Despite the progress in Multimodal Large Language Models (MLLMs), reasoning performance…

人工智能 · 计算机科学 2026-04-07 Xiaoyu Chen , Lu Dai , Hanqing Wang , Zhuoyu Li , Wenbin Dai , Yanzong Zheng , Zhenggang Xia , Junyong Lin , Hui Xiong

Spatial reasoning is central to navigation and robotics, yet measuring model capabilities on these tasks remains difficult. Existing benchmarks evaluate models in a one-shot setting, requiring full solution generation in a single response,…

人工智能 · 计算机科学 2026-04-13 Lars Benedikt Kaesberg , Tianyu Yang , Niklas Bauer , Terry Ruas , Jan Philip Wahle , Bela Gipp

Recent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, extending these successes to vision-language tasks remains…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Minheng Ni , Zhengyuan Yang , Linjie Li , Chung-Ching Lin , Kevin Lin , Wangmeng Zuo , Lijuan Wang

Multimodal Large Language Models (MLLMs) based agents have demonstrated remarkable potential in autonomous web navigation. However, handling long-horizon tasks remains a critical bottleneck. Prevailing strategies often rely heavily on…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Dawei Yan , Haokui Zhang , Guangda Huzhang , Yang Li , Yibo Wang , Qing-Guo Chen , Zhao Xu , Weihua Luo , Ying Li , Wei Dong , Chunhua Shen

Existing Multimodal Large Language Models (MLLMs) struggle with 3D spatial reasoning, as they fail to construct structured abstractions of the 3D environment depicted in video inputs. To bridge this gap, drawing inspiration from cognitive…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Jiacheng Hua , Yishu Yin , Yuhang Wu , Tai Wang , Yifei Huang , Miao Liu

Embodied intelligence, a grand challenge in artificial intelligence, is fundamentally constrained by the limited spatial understanding and reasoning capabilities of current models. Prevailing efforts to address this through enhancing…

In this paper, we address the challenging task of multimodal reasoning by incorporating the notion of ``slow thinking'' into multimodal large language models (MLLMs). Our core idea is that models can learn to adaptively use different levels…

When presented with questions involving visual thinking, humans naturally switch reasoning modalities, often forming mental images or drawing visual aids. Large language models have shown promising results in arithmetic and symbolic…

计算与语言 · 计算机科学 2024-06-21 Sachit Menon , Richard Zemel , Carl Vondrick

Document understanding with multimodal large language models (MLLMs) requires not only accurate answers but also explicit, evidence-grounded reasoning, especially in high-stakes scenarios. However, current document MLLMs still fall short of…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Yuchuan Wu , Minghan Zhuo , Teng Fu , Mengyang Zhao , Bin Li , Xiangyang Xue

Reliable spatial reasoning remains a core bottleneck for vision-language models (VLMs). Existing mainstream training paradigms for spatial reasoning largely rely on outcome alignment or process imitation, lacking explicit constraints on the…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Jiangyang Li , Cong Wan , Changjie Wu , Songlin Dong , Lingjun Zhang , Linzhe Shi , Xu Wang , Zhiheng Ma , Hang Zhang , Mu Xu , Yihong Gong

Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existing studies primarily focus on logical reasoning and visual…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Dingning Liu , Cheng Wang , Peng Gao , Renrui Zhang , Xinzhu Ma , Yuan Meng , Zhihui Wang

Multi-step spatial reasoning entails understanding and reasoning about spatial relationships across multiple sequential steps, which is crucial for tackling complex real-world applications, such as robotic manipulation, autonomous…

人工智能 · 计算机科学 2025-06-23 Kexian Tang , Junyao Gao , Yanhong Zeng , Haodong Duan , Yanan Sun , Zhening Xing , Wenran Liu , Kaifeng Lyu , Kai Chen

Table reasoning requires models to jointly perform semantic understanding and precise numerical operations. Most existing methods rely on a single-turn reasoning paradigm over tables which suffers from context overflow and weak numerical…

计算与语言 · 计算机科学 2026-03-11 Mingyue Cheng , Shuo Yu , Chuang Jiang , Xiaoyu Tao , Qingyang Mao , Jie Ouyang , Qi Liu , Enhong Chen

We introduce RoboBrain 2.0, our latest generation of embodied vision-language foundation models, designed to unify perception, reasoning, and planning for complex embodied tasks in physical environments. It comes in two variants: a…

Multimodal Large Language Models (MLLMs) have made remarkable progress in multimodal perception and reasoning by bridging vision and language. However, most existing MLLMs perform reasoning primarily with textual CoT, which limits their…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Jintao Tong , Shilin Yan , Hongwei Xue , Xiaojun Tang , Kunyu Shi , Guannan Zhang , Ruixuan Li , Yixiong Zou

Spatial intelligence requires multimodal large language models (MLLMs) to move beyond single-view perception and reason consistently about objects, visibility, geometry, and interactions across multiple viewpoints. However, progress in…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Wei Wang , Yuqian Yuan , Tianwei Lin , Wenqiao Zhang , Siliang Tang , Jun Xiao , Yueting Zhuang

Spatial reasoning in large-scale 3D environments remains challenging for current vision-language models, which are typically constrained to room-scale scenarios. We introduce H$^2$U3D (Holistic House Understanding in 3D), a 3D visual…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Hongpei Zheng , Shijie Li , Yanran Li , Hujun Yin

Multi-view visual reasoning is essential for intelligent systems that must understand complex environments from sparse and discrete viewpoints, yet existing research has largely focused on single-image or temporally dense video settings. In…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Fucai Ke , Zhixi Cai , Boying Li , Long Chen , Beibei Lin , Weiqing Wang , Pari Delir Haghighi , Gholamreza Haffari , Hamid Rezatofighi

While Multi-Modal Large Language Models (MLLMs) demonstrate impressive capabilities in general reasoning, their embodied spatial intelligence remains hampered by a "Cartesian Illusion" - a reliance on text-based probability distributions…

人工智能 · 计算机科学 2026-05-19 Yajing Zhou , Xiangyu Kong