English
Related papers

Related papers: GeoAlign: Geometric Feature Realignment for MLLM S…

200 papers

Visual Language Models (VLMs) have increasingly become the main paradigm for understanding indoor scenes, but they still struggle with metric and spatial reasoning. Current approaches rely on end-to-end video understanding or large-scale…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Fernando Ropero , Erkin Turkoz , Daniel Matos , Junqing Du , Antonio Ruiz , Yanfeng Zhang , Lu Liu , Mingwei Sun , Yongliang Wang

Large Multimodal Models (LMMs) often struggle with geometric reasoning due to visual hallucinations and a lack of mathematically precise Chain-of-Thought (CoT) data. To address this, we propose the GeoSym Engine, an automated and scalable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jinhao Jing , Zheng Ma , Jinwei Liang , Qiannian Zhao , Shawn Chen , Jing Yang , Por Lip Yee , Prayag Tiwari , Jingjing Bai , Benyou Wang , Lewei Lu , Zhan Su

Improving embodied reasoning in multimodal-large-language models (MLLMs) is essential for building vision-language-action models (VLAs) on top of them to readily translate multimodal understanding into low-level actions. Accordingly, recent…

Artificial Intelligence · Computer Science 2026-03-24 Dongyoung Kim , Sumin Park , Woomin Song , Seungku Kim , Taeyoung Kim , Huiwon Jang , Jinwoo Shin , Jaehyung Kim , Younggyo Seo

4D spatial intelligence involves perceiving and processing how objects move or change over time. Humans naturally possess 4D spatial intelligence, supporting a broad spectrum of spatial reasoning abilities. To what extent can Multimodal…

Large Vision-Language Models (LVLMs) exhibit powerful reasoning capabilities but suffer sophisticated jailbreak vulnerabilities. Fundamentally, aligning LVLMs is not just a safety challenge but a problem of economic efficiency. Current…

Artificial Intelligence · Computer Science 2026-03-17 Ruoxi Cheng , Haoxuan Ma , Teng Ma , Hongyi Zhang

Mathematical reasoning is a hallmark of human intelligence, requiring logical deduction, symbolic manipulation, and abstract thinking. Recent multimodal large language models (MLLMs) have demonstrated strong performance on geometry problems…

Computation and Language · Computer Science 2026-05-26 Yingji Zhang , Yong Dai , André Freitas

While Multimodal Large Language Models (MLLMs) have achieved impressive performance on semantic tasks, their spatial intelligence--crucial for robust and grounded AI systems--remains underdeveloped. Existing benchmarks fall short of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Mingrui Wu , Zhaozhi Wang , Fangjinhua Wang , Jiaolong Yang , Marc Pollefeys , Tong Zhang

Multimodal large language models (MLLMs) have made rapid progress in recent years, yet continue to struggle with low-level visual perception (LLVP) -- particularly the ability to accurately describe the geometric details of an image. This…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Jiarui Zhang , Ollie Liu , Tianyu Yu , Jinyi Hu , Willie Neiswanger

Multimodal large language models (MLLMs) have achieved strong performance on perception-oriented tasks, yet their ability to perform mathematical spatial reasoning, defined as the capacity to parse and manipulate two- and three-dimensional…

Despite strong performance on vision-language tasks, Multimodal Large Language Models (MLLMs) struggle with mathematical problem-solving, with both open-source and state-of-the-art models falling short of human performance on visual-math…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 William Rudman , Michal Golovanevsky , Amir Bar , Vedant Palit , Yann LeCun , Carsten Eickhoff , Ritambhara Singh

Geometry problem-solving remains a significant challenge for Large Multimodal Models (LMMs), requiring not only global shape recognition but also attention to intricate local relationships related to geometric theory. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Linger Deng , Yuliang Liu , Wenwen Yu , Zujia Zhang , Jianzhong Ju , Zhenbo Luo , Xiang Bai

Multi-modal Large Language Models (MLLMs) exhibit impressive problem-solving abilities in various domains, but their visual comprehension and abstract reasoning skills remain under-evaluated. To this end, we present PolyMATH, a challenging…

Artificial Intelligence · Computer Science 2026-05-12 Himanshu Gupta , Shreyas Verma , Ujjwala Anantheswaran , Kevin Scaria , Mihir Parmar , Swaroop Mishra , Chitta Baral

Multi-modal large language models (MLLMs), such as GPT-4o, excel at integrating text and visual data but face systematic challenges when interpreting ambiguous or incomplete visual stimuli. This study leverages statistical modeling to…

Machine Learning · Computer Science 2024-12-09 Ching-Yi Wang

Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared…

The development of 3D Vision-Language Models (VLMs), crucial for applications in robotics, autonomous driving, and augmented reality, is severely constrained by the scarcity of paired 3D-text data. Existing methods rely solely on next-token…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yuanhao Su , Shaofeng Zhang , Xiaosong Jia , Qi Fan

The integration of Large Language Models (LLMs) into Geographic Information Systems (GIS) marks a paradigm shift toward autonomous spatial analysis. However, evaluating these LLM-based agents remains challenging due to the complex,…

Artificial Intelligence · Computer Science 2026-04-16 Bo Yu , Cheng Yang , Dongyang Hou , Chengfu Liu , Jiayao Liu , Chi Wang , Zhiming Zhang , Haifeng Li , Wentao Yang

Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools or image generators. However, existing methods usually follow an output-as-input latent…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Yanting Miao , Yutao Sun , Dexin Wang , Mengyu Zhou , Pascal Poupart , Lei Lv , Qi Zhao , Li Wang , Hao Li , Xiaoxi Jiang , Guanjun Jiang

Multimodal large language models (MLLMs) are proficient in perception and instruction-following, but they still struggle with spatial reasoning: the ability to mentally track and manipulate objects across multiple views and over time.…

Artificial Intelligence · Computer Science 2025-12-30 Ryan Spencer , Roey Yaari , Ritvik Vemavarapu , Joyce Yang , Steven Ngo , Utkarsh Sharma

Modern monocular 3D reconstruction methods and vision-language models (VLMs) demonstrate impressive results on standard benchmarks, yet recent works cast doubt on their true understanding of geometric properties. We introduce GOQ, a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Mateusz Michalkiewicz , Anekha Sokhal , Tadeusz Michalkiewicz , Piotr Pawlikowski , Mahsa Baktashmotlagh , Varun Jampani , Guha Balakrishnan

Interior design is a requirements-to-visual-plan generation process that must simultaneously satisfy verifiable spatial feasibility and comparative aesthetic preferences. While recent multimodal large language models (MLLMs) offer a unified…

Multimedia · Computer Science 2026-03-17 Yuxuan Yang , Xiaotong Mao , Jingyao Wang , Fuchun Sun
‹ Prev 1 4 5 6 7 8 10 Next ›