English
Related papers

Related papers: Geoint-R1: Formalizing Multimodal Geometric Reason…

200 papers

A key frontier for Multimodal Large Language Models (MLLMs) is the ability to perform deep mathematical and spatial reasoning directly from images, moving beyond their established success in semantic description. Mathematical surface plots…

Artificial Intelligence · Computer Science 2025-09-10 Nilay Pande , Sahiti Yerramilli , Jayant Sravan Tamarapalli , Rynaa Grover

Despite strong performance on vision-language tasks, Multimodal Large Language Models (MLLMs) struggle with mathematical problem-solving, with both open-source and state-of-the-art models falling short of human performance on visual-math…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 William Rudman , Michal Golovanevsky , Amir Bar , Vedant Palit , Yann LeCun , Carsten Eickhoff , Ritambhara Singh

Remote sensing has become critical for understanding environmental dynamics, urban planning, and disaster management. However, traditional remote sensing workflows often rely on explicit segmentation or detection methods, which struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Kaiyu Li , Zepeng Xin , Li Pang , Chao Pang , Yupeng Deng , Jing Yao , Guisong Xia , Deyu Meng , Zhi Wang , Xiangyong Cao

Deeply understanding sports requires an intricate blend of fine-grained visual perception and rule-based reasoning - a challenge that pushes the limits of current multimodal models. To succeed, models must master three critical…

Unified multimodal models aim to jointly enable visual understanding and generation, yet current benchmarks rarely examine their true integration. Existing evaluations either treat the two abilities in isolation or overlook tasks that…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Kai Zou , Ziqi Huang , Yuhao Dong , Shulin Tian , Dian Zheng , Hongbo Liu , Jingwen He , Bin Liu , Yu Qiao , Ziwei Liu

While rule-based reinforcement learning has recently catalyzed explicit reasoning in multimodal models, tactile reasoning remains largely underexplored. Existing tactile-language models primarily rely on supervised or contrastive…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Yingxin Lai , Yafei Zhou , Fucai Zhu , Siyu Zhu , Weihao Yuan

As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdisciplinary, multimodal, and application-driven nature. However, existing materials…

Artificial Intelligence · Computer Science 2026-05-29 Wanhao Liu , Jiaqing Xie , Qian Tan , Weida Wang , Jue Wang , Ran Sun , Zhuo Yang , Wanli Ouyang , Lei Bai , Tianfan Fu , Lu Chen , Xin Chen , Yuqiang Li

Accurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual perception with natural language, current medical-grounding…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Zhonghao Yan , Muxi Diao , Yuxuan Yang , Ruoyan Jing , Jiayuan Xu , Kaizhou Zhang , Lele Yang , Yanxi Liu , Kongming Liang , Zhanyu Ma

The research in AI-based formal mathematical reasoning has shown an unstoppable growth trend. These studies have excelled in mathematical competitions like IMO and have made significant progress. This paper focuses on formal verification,…

Artificial Intelligence · Computer Science 2025-06-10 Jialun Cao , Yaojie Lu , Meiziniu Li , Haoyang Ma , Haokun Li , Mengda He , Cheng Wen , Le Sun , Hongyu Zhang , Shengchao Qin , Shing-Chi Cheung , Cong Tian

Recent advancements in Large Vision-Language Models (LVLMs) have significantly enhanced their ability to integrate visual and linguistic information, achieving near-human proficiency in tasks like object recognition, captioning, and visual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Zhikai Wang , Jiashuo Sun , Wenqi Zhang , Zhiqiang Hu , Xin Li , Fan Wang , Deli Zhao

This paper proposes a novel approach to analyzing multi-hop reasoning in language models through Hamiltonian mechanics. We map reasoning chains in embedding spaces to Hamiltonian systems, defining a function that balances reasoning…

Artificial Intelligence · Computer Science 2025-03-11 Javier Marin

Numerous theorems, such as those in geometry, are often presented in multimodal forms (e.g., diagrams). Humans benefit from visual reasoning in such settings, using diagrams to gain intuition and guide the proof process. Modern Multimodal…

Computation and Language · Computer Science 2025-06-09 Zhitao He , Zongwei Lyu , Dazhong Chen , Dadi Guo , Yi R. Fung

While Vision-Language Models (VLMs) have significantly advanced remote sensing interpretation, enabling them to perform complex, step-by-step reasoning remains highly challenging. Recent efforts to introduce Chain-of-Thought (CoT) reasoning…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Lang Sun , Ronghao Fu , Zhuoran Duan , Haoran Liu , Xueyan Liu , Bo Yang

Large Multimodal Models (LMMs) are increasingly applied to scientific research, yet it remains unclear whether they can reliably understand and reason over the multimodal complexity of papers. A central challenge lies in detecting and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Lukas Selch , Yufang Hou , M. Jehanzeb Mirza , Sivan Doveh , James Glass , Rogerio Feris , Wei Lin

The ability to process information from multiple modalities and to reason through it step-by-step remains a critical challenge in advancing artificial intelligence. However, existing reasoning benchmarks focus on text-only reasoning, or…

Artificial Intelligence · Computer Science 2025-07-01 Yulun Jiang , Yekun Chai , Maria Brbić , Michael Moor

Modern monocular 3D reconstruction methods and vision-language models (VLMs) demonstrate impressive results on standard benchmarks, yet recent works cast doubt on their true understanding of geometric properties. We introduce GOQ, a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Mateusz Michalkiewicz , Anekha Sokhal , Tadeusz Michalkiewicz , Piotr Pawlikowski , Mahsa Baktashmotlagh , Varun Jampani , Guha Balakrishnan

Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity. We introduce MathNet, a high-quality,…

Artificial Intelligence · Computer Science 2026-04-21 Shaden Alshammari , Kevin Wen , Abrar Zainal , Mark Hamilton , Navid Safaei , Sultan Albarakati , William T. Freeman , Antonio Torralba

Lithology classification in well logs is a fundamental geoscience data mining task that aims to infer rock types from multi dimensional geophysical sequences. Despite recent progress, existing approaches typically formulate the problem as a…

Artificial Intelligence · Computer Science 2026-04-28 Yitong Zhou , Mingyue Cheng , Jiahao Wang , Qingyang Mao , Qi Liu

The rise of Multimodal Large Language Models (MLLMs), renowned for their advanced instruction-following and reasoning capabilities, has significantly propelled the field of visual reasoning. However, due to limitations in their image…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Jiaxing Chen , Yuxuan Liu , Dehu Li , Xiang An , Weimo Deng , Ziyong Feng , Yongle Zhao , Yin Xie

Large Multimodal Models have achieved remarkable progress in integrating vision and language, enabling strong performance across perception, reasoning, and domain-specific tasks. However, their capacity to reason over multiple, visually…

Artificial Intelligence · Computer Science 2026-03-09 Can Li , Ying Liu , Ting Zhang , Mei Wang , Hua Huang