中文
相关论文

相关论文: How Far Are Large Multimodal Models from Human-Lev…

200 篇论文

Large language models (LLMs) and large multimodal models (LMMs) have achieved unprecedented breakthrough, showcasing remarkable capabilities in natural language understanding, generation, and complex reasoning. This transformative potential…

机器学习 · 计算机科学 2025-10-24 Hyun Jong Yang , Hyunsoo Kim , Hyeonho Noh , Seungnyun Kim , Byonghyo Shim

Large Language Models (LLMs) have demonstrated potential in Vision-and-Language Navigation (VLN) tasks, yet current applications face challenges. While LLMs excel in general conversation scenarios, they struggle with specialized navigation…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Yunzhe Xu , Yiyuan Pan , Zhe Liu , Hesheng Wang

While Multimodal Large Language Models (MLLMs) have achieved impressive performance on semantic tasks, their spatial intelligence--crucial for robust and grounded AI systems--remains underdeveloped. Existing benchmarks fall short of…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Mingrui Wu , Zhaozhi Wang , Fangjinhua Wang , Jiaolong Yang , Marc Pollefeys , Tong Zhang

Spatial reasoning is a fundamental capability for embodied intelligence, especially for fine-grained manipulation tasks such as robotic assembly. While recent vision-language models (VLMs) exhibit preliminary spatial awareness, they largely…

机器人学 · 计算机科学 2026-04-13 Zhi Jing , Jinbin Qiao , Ouyang Lu , Jicong Ao , Shuang Qiu , Yu-Gang Jiang , Chenjia Bai

Despite the ubiquity of large language models (LLMs) in AI research, the question of embodiment in LLMs remains underexplored, distinguishing them from embodied systems in robotics where sensory perception directly informs physical action.…

计算与语言 · 计算机科学 2024-05-28 Philipp Wicke , Lennart Wachowiak

In the evolving landscape of transportation systems, integrating Large Language Models (LLMs) offers a promising frontier for advancing intelligent decision-making across various applications. This paper introduces a novel 3-dimensional…

机器学习 · 计算机科学 2024-12-17 Dexter Le , Aybars Yunusoglu , Karn Tiwari , Murat Isik , I. Can Dikmen

The advent of Large Multimodal Models (LMMs) offers a promising technology to tackle the limitations of modular design in autonomous driving, which often falters in open-world scenarios requiring sustained environmental understanding and…

机器人学 · 计算机科学 2026-01-21 Long Zhang , Yuchen Xia , Bingqing Wei , Zhen Liu , Shiwen Mao , Zhu Han , Mohsen Guizani

Spatial reasoning is a fundamental capability of multimodal large language models (MLLMs), yet their performance in open aerial environments remains underexplored. In this work, we present Open3D-VQA, a novel benchmark for evaluating MLLMs'…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Weichen Zhang , Zile Zhou , Xin Zeng , Xuchen Liu , Jianjie Fang , Chen Gao , Yong Li , Jinqiang Cui , Xinlei Chen , Xiao-Ping Zhang

Multimodal Large Language Models (MLLMs) have demonstrated a wide range of capabilities across many domains, including Embodied AI. In this work, we study how to best ground a MLLM into different embodiments and their associated action…

机器学习 · 计算机科学 2024-12-10 Andrew Szot , Bogdan Mazoure , Harsh Agrawal , Devon Hjelm , Zsolt Kira , Alexander Toshev

Multimodal Large Language Models (MLLMs) have demonstrated extraordinary progress in bridging textual and visual inputs. However, MLLMs still face challenges in situated physical and social interactions in sensorally rich, multimodal and…

神经元与认知 · 定量生物学 2025-10-17 Akila Kadambi , Lisa Aziz-Zadeh , Antonio Damasio , Marco Iacoboni , Srini Narayanan

Multimodal large language models (MLLMs) have demonstrated powerful capabilities in general spatial understanding and reasoning. However, their fine-grained spatial understanding and reasoning capabilities in complex urban scenarios have…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Jun Zhang , Jie Feng , Long Chen , Junhui Wang , Zhicheng Liu , Depeng Jin , Yong Li

While Multimodal Large Language Models (MLLMs) have exhibited remarkable general intelligence across diverse domains, their potential in low-altitude applications dominated by Unmanned Aerial Vehicles (UAVs) remains largely underexplored.…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Shiqi Dai , Zizhi Ma , Zhicong Luo , Xuesong Yang , Yibin Huang , Wanyue Zhang , Chi Chen , Zonghao Guo , Wang Xu , Yufei Sun , Maosong Sun

Humans possess the visual-spatial intelligence to remember spaces from sequential visual observations. However, can Multimodal Large Language Models (MLLMs) trained on million-scale video datasets also ``think in space'' from videos? We…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Jihan Yang , Shusheng Yang , Anjali W. Gupta , Rilyn Han , Li Fei-Fei , Saining Xie

We explore the human motion knowledge of Large Language Models (LLMs) through 3D avatar control. Given a motion instruction, we prompt LLMs to first generate a high-level movement plan with consecutive steps (High-level Planning), then…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Kunhang Li , Jason Naradowsky , Yansong Feng , Yusuke Miyao

Spatial cognition is fundamental to real-world multimodal intelligence, allowing models to effectively interact with the physical environment. While multimodal large language models (MLLMs) have made significant strides, existing benchmarks…

人工智能 · 计算机科学 2026-05-08 Peiran Xu , Sudong Wang , Yao Zhu , Jianing Li , Gege Qi , Yunjian Zhang

Spatial reasoning, which requires ability to perceive and manipulate spatial relationships in the 3D world, is a fundamental aspect of human intelligence, yet remains a persistent challenge for Multimodal large language models (MLLMs).…

人工智能 · 计算机科学 2025-11-21 Weichen Liu , Qiyao Xue , Haoming Wang , Xiangyu Yin , Boyuan Yang , Wei Gao

Multi-step spatial reasoning entails understanding and reasoning about spatial relationships across multiple sequential steps, which is crucial for tackling complex real-world applications, such as robotic manipulation, autonomous…

人工智能 · 计算机科学 2025-06-23 Kexian Tang , Junyao Gao , Yanhong Zeng , Haodong Duan , Yanan Sun , Zhening Xing , Wenran Liu , Kaifeng Lyu , Kai Chen

Although large multimodal models (LMMs) have demonstrated remarkable capabilities in visual scene interpretation and reasoning, their capacity for complex and precise 3-dimensional spatial reasoning remains uncertain. Existing benchmarks…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Xingrui Wang , Wufei Ma , Tiezheng Zhang , Celso M de Melo , Jieneng Chen , Alan Yuille

Large language models (LLMs) have undergone significant expansion and have been increasingly integrated across various domains. Notably, in the realm of robot task planning, LLMs harness their advanced reasoning and language comprehension…

Understanding how humans conceptualize and categorize natural objects offers critical insights into perception and cognition. With the advent of Large Language Models (LLMs), a key question arises: can these models develop human-like object…