中文
相关论文

相关论文: GRASP: A novel benchmark for evaluating language G…

200 篇论文

Geospatial pixel reasoning aims to generate segmentation masks in remote sensing imagery directly from natural-language instructions. Most existing approaches follow a paradigm that fine-tunes multimodal large language models under…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Chengjie Jiang , Yunqi Zhou , Jiafeng Yan , Jing Li , Jiayang Li , Yue Zhou , Hongjie He , Jonathan Li

The fusion of Large Language Models (LLMs) and robotic systems has led to a transformative paradigm in the robotic field, offering unparalleled capabilities not only in the communication domain but also in skills like multimodal input…

机器人学 · 计算机科学 2025-02-18 Sara Incao , Carlo Mazzola , Giulia Belgiovine , Alessandra Sciutti

In current benchmarks for evaluating large language models (LLMs), there are issues such as evaluation content restriction, untimely updates, and lack of optimization guidance. In this paper, we propose a new paradigm for the measurement of…

计算与语言 · 计算机科学 2024-07-11 Jin Liu , Qingquan Li , Wenlong Du

Key to tasks that require reasoning about natural language in visual contexts is grounding words and phrases to image regions. However, observing this grounding in contemporary models is complex, even if it is generally expected to take…

计算与语言 · 计算机科学 2024-06-03 Noriyuki Kojima , Hadar Averbuch-Elor , Yoav Artzi

Multimodal Large Language Models (MLLMs) rely on strong linguistic reasoning inherited from their base language models. However, multimodal instruction fine-tuning paradoxically degrades this text's reasoning capability, undermining…

Humans build viewpoint-independent cognitive maps through navigation, enabling intuitive reasoning about object permanence and spatial relations. We argue that multimodal large language models (MLLMs), despite extensive video training, lack…

机器学习 · 计算机科学 2025-12-02 Jacob Thompson , Emiliano Garcia-Lopez , Yonatan Bisk

Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-answering (VQA) datasets may overfit dataset-specific biases, while integrating specialized…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Chun-Hsiao Yeh , Shengyi Qian , Manchen Wang , Yi Ma , Joseph Tighe , Fanyi Xiao

Recent studies have demonstrated that many layers are functionally redundant in large language models (LLMs), enabling model compression by removing these layers to reduce inference cost. While such approaches can improve efficiency,…

计算与语言 · 计算机科学 2026-02-24 Kainan Liu , Yong Zhang , Ning Cheng , Zhitao Li , Shaojun Wang , Jing Xiao

Several machine learning methods aim to learn or reason about complex physical systems. A common first-step towards reasoning is to infer system parameters from observations of its behavior. In this paper, we investigate the performance of…

计算与语言 · 计算机科学 2024-02-07 Sean Memery , Mirella Lapata , Kartic Subr

Modern neural language models (LMs) are powerful tools for modeling human sentence production and comprehension, and their internal representations are remarkably well-aligned with representations of language in the human brain. But to…

计算与语言 · 计算机科学 2024-03-27 Chengxu Zhuang , Evelina Fedorenko , Jacob Andreas

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with precise pixel-level…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Shehan Munasinghe , Hanan Gani , Wenqi Zhu , Jiale Cao , Eric Xing , Fahad Shahbaz Khan , Salman Khan

Large language models (LLMs) and vision-language models (VLMs) have demonstrated remarkable performance across a wide range of tasks and domains. Despite this promise, spatial understanding and reasoning -- a fundamental component of human…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Jiayu Wang , Yifei Ming , Zhenmei Shi , Vibhav Vineet , Xin Wang , Yixuan Li , Neel Joshi

Recent advances in Vision-Language Models (VLMs) have shown promising capabilities in interpreting visualized graph data, offering a new perspective for graph-structured reasoning beyond traditional Graph Neural Networks (GNNs). However,…

人工智能 · 计算机科学 2026-04-27 Qihang Ai , Ruizhou Li , Menghui Wang , Haiyun Jiang

Most multimodal large language models (MLLMs) learn language-to-object grounding through causal language modeling where grounded objects are captured by bounding boxes as sequences of location tokens. This paradigm lacks pixel-level…

计算机视觉与模式识别 · 计算机科学 2024-04-17 Yichi Zhang , Ziqiao Ma , Xiaofeng Gao , Suhaila Shakiah , Qiaozi Gao , Joyce Chai

The integration of language and 3D perception is crucial for embodied agents and robots that comprehend and interact with the physical world. While large language models (LLMs) have demonstrated impressive language understanding and…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Jianing Yang , Xuweiyi Chen , Nikhil Madaan , Madhavan Iyengar , Shengyi Qian , David F. Fouhey , Joyce Chai

Current state of the art measures like BLEU, CIDEr, VQA score, SigLIP-2 and CLIPScore are often unable to capture semantic or structural accuracy, especially for domain-specific or context-dependent scenarios. For this, this paper proposes…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Kishor Datta Gupta , Marufa Kamal , Md. Mahfuzur Rahman , Fahad Rahman , Mohd Ariful Haque , Sunzida Siddique

This paper introduces the TempVS benchmark, which focuses on temporal grounding and reasoning capabilities of Multimodal Large Language Models (MLLMs) in image sequences. TempVS consists of three main tests (i.e., event relation inference,…

计算与语言 · 计算机科学 2025-06-13 Yingjin Song , Yupei Du , Denis Paperno , Albert Gatt

In recent years, Multimodal Large Language Models (MLLMs) have achieved remarkable progress on a wide range of multimodal benchmarks. Despite these advances, most existing benchmarks mainly focus on single-image or multi-image…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Bingli Wang , Huanze Tang , Haijun Lv , Zhishan Lin , Lixin Gu , Lei Feng , Qipeng Guo , Kai Chen

Large-scale, pre-trained language models (LMs) have achieved human-level performance on a breadth of language understanding tasks. However, evaluations only based on end task performance shed little light on machines' true ability in…

计算与语言 · 计算机科学 2022-05-11 Shane Storks , Qiaozi Gao , Yichi Zhang , Joyce Chai

Large Language Models (LLMs) show potential for enhancing robotic path planning. This paper assesses visual input's utility for multimodal LLMs in such tasks via a comprehensive benchmark. We evaluated 15 multimodal LLMs on generating valid…

机器人学 · 计算机科学 2025-07-17 Jacinto Colan , Ana Davila , Yasuhisa Hasegawa