English
Related papers

Related papers: GeoX: Mastering Geospatial Reasoning Through Self-…

200 papers

Humans are able to identify a referred visual object in a complex scene via a few rounds of natural language communications. Success communication requires both parties to engage and learn to adapt for each other. In this paper, we…

Artificial Intelligence · Computer Science 2017-12-05 Yan Zhu , Shaoting Zhang , Dimitris Metaxas

Multimodal large language models (MLLMs) have achieved strong performance on perception-oriented tasks, yet their ability to perform mathematical spatial reasoning, defined as the capacity to parse and manipulate two- and three-dimensional…

Scalable sensor simulation is an important yet challenging open problem for safety-critical domains such as self-driving. Current works in image simulation either fail to be photorealistic or do not model the 3D environment and the dynamic…

Computer Vision and Pattern Recognition · Computer Science 2021-05-18 Yun Chen , Frieda Rong , Shivam Duggal , Shenlong Wang , Xinchen Yan , Sivabalan Manivasagam , Shangjie Xue , Ersin Yumer , Raquel Urtasun

In this paper, we claim that 3D visual grounding is the cornerstone of spatial reasoning and introduce the Grounded-Spatial Reasoner (GS-Reasoner) to explore the effective spatial representations that bridge the gap between them. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Yiming Chen , Zekun Qi , Wenyao Zhang , Xin Jin , Li Zhang , Peidong Liu

Maps are powerful carriers of structured and contextual knowledge, encompassing geography, demographics, infrastructure, and environmental patterns. Reasoning over such knowledge requires models to integrate spatial relationships, visual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Sharat Bhat , Harshita Khandelwal , Tushar Kataria , Vivek Gupta

Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language models. Despite promising performance, recent multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yian Li , Yang Jiao , Bin Zhu , Tianwen Qian , Shaoxiang Chen , Jingjing Chen , Yu-Gang Jiang

AI-driven geometric problem solving is a complex vision-language task that requires accurate diagram interpretation, mathematical reasoning, and robust cross-modal grounding. A foundational yet underexplored capability for this task is the…

Machine Learning · Computer Science 2025-09-26 Bing Liu , Wenqiang Yv , Xuzheng Yang , Shichang Wang , Junzhuo Liu , Peng Wang , Guoqing Wang , Yang Yang , Heng Tao Shen

Our ability to know when to trust the decisions made by machine learning systems has not kept up with the staggering improvements in their performance, limiting their applicability in high-stakes domains. We introduce Prover-Verifier Games…

Machine Learning · Computer Science 2021-08-30 Cem Anil , Guodong Zhang , Yuhuai Wu , Roger Grosse

While the situation has improved for text-only models, it again seems to be the case currently that multimodal (text and image) models develop faster than ways to evaluate them. In this paper, we bring a recently developed evaluation…

Computation and Language · Computer Science 2024-12-12 Sherzod Hakimov , Yerkezhan Abdullayeva , Kushal Koshti , Antonia Schmidt , Yan Weiser , Anne Beyer , David Schlangen

Graphical user interface (GUI) grounding is a fundamental task for building GUI agents. However, general vision-language models (VLMs) struggle with this task due to a lack of specific optimization. We identify a key gap in this paper:…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Weiming Li , Yan Shao , Jing Yang , Yujing Lu , Ling Zhong , Yuhan Wang , Manni Duan

Automatic math problem solving has recently attracted increasing attention as a long-standing AI benchmark. In this paper, we focus on solving geometric problems, which requires a comprehensive understanding of textual descriptions, visual…

Artificial Intelligence · Computer Science 2022-01-12 Jiaqi Chen , Jianheng Tang , Jinghui Qin , Xiaodan Liang , Lingbo Liu , Eric P. Xing , Liang Lin

A common approach to solving physical reasoning tasks is to train a value learner on example tasks. A limitation of such an approach is that it requires learning about object dynamics solely from reward values assigned to the final state of…

Artificial Intelligence · Computer Science 2021-09-03 Eltayeb Ahmed , Anton Bakhtin , Laurens van der Maaten , Rohit Girdhar

Reasoning is a fundamental capability of large language models (LLMs), enabling them to comprehend, analyze, and solve complex problems. In this paper, we introduce TextGames, an innovative benchmark specifically crafted to assess LLMs…

Computation and Language · Computer Science 2025-02-26 Frederikus Hudi , Genta Indra Winata , Ruochen Zhang , Alham Fikri Aji

Geometric reasoning inherently requires "thinking with constructions" -- the dynamic manipulation of visual aids to bridge the gap between problem conditions and solutions. However, existing Multimodal Large Language Models (MLLMs) are…

Artificial Intelligence · Computer Science 2026-03-20 Haokun Zhao , Wanshi Xu , Haidong Yuan , Songjun Cao , Long Ma , Yanghua Xiao

As Earth science enters the era of big data, artificial intelligence (AI) not only offers great potential for solving geoscience problems, but also plays a critical role in accelerating the understanding of the complex, interactive, and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Jin-Jian Xu , Hao Zhang , Chao-Sheng Tang , Lin Li , Bin Shi

Spatial intelligence spans a rich suite of abilities, including visualising and transforming shapes, mentally rotating objects, judging relational positions and containment, and estimating numerosity. However, it still remains a critical…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Shijie Lian , Changti Wu , Laurence Tianruo Yang , Hang Yuan , Bin Yu , Lei Zhang , Kai Chen

Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such representations directly from unposed multi-view images…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Bo Zhou , Qiuxia Lai , Zeren Sun , Xiangbo Shu , Yazhou Yao , Wenguan Wang

The growing interest in embodied agents increases the demand for spatiotemporal video understanding, yet existing benchmarks largely emphasize extractive reasoning, where answers can be explicitly presented within spatiotemporal events. It…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Seunghwan Bang , Hwanjun Song

The human-like automatic deductive reasoning has always been one of the most challenging open problems in the interdiscipline of mathematics and artificial intelligence. This paper is the third in a series of our works. We built a…

Artificial Intelligence · Computer Science 2024-02-16 Jia Zou , Xiaokai Zhang , Yiming He , Na Zhu , Tuo Leng

In order for an autonomous robot to efficiently explore an unknown environment, it must account for uncertainty in sensor measurements, hazard assessment, localization, and motion execution. Making decisions for maximal reward in a…