English
Related papers

Related papers: GeoVista: Web-Augmented Agentic Visual Reasoning f…

200 papers

Video agentic models have advanced challenging video-language tasks. However, most agentic approaches still heavily rely on greedy parsing over densely sampled video frames, resulting in high computational cost. We present VideoSeek, a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Jingyang Lin , Jialian Wu , Jiang Liu , Ximeng Sun , Ze Wang , Xiaodong Yu , Jiebo Luo , Zicheng Liu , Emad Barsoum

Three-dimensional geospatial analysis is critical for applications in urban planning, climate adaptation, and environmental assessment. However, current methodologies depend on costly, specialized sensors, such as LiDAR and multispectral…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Mai Tsujimoto , Junjue Wang , Weihao Xuan , Naoto Yokoya

Reinforcement learning for agentic multimodal models often suffers from interaction collapse, where models learn to reduce tool usage and multi-turn reasoning, limiting the benefits of agentic behavior. We introduce PyVision-RL, a…

Artificial Intelligence · Computer Science 2026-02-25 Shitian Zhao , Shaoheng Lin , Ming Li , Haoquan Zhang , Wenshuo Peng , Kaipeng Zhang , Chen Wei

Recent advances in reinforcement learning (RL) have delivered strong reasoning capabilities in natural image domains, yet their potential for Earth Observation (EO) remains largely unexplored. EO tasks introduce unique challenges, spanning…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Mustansar Fiaz , Hiyam Debary , Paolo Fraccaro , Danda Paudel , Luc Van Gool , Fahad Khan , Salman Khan

Graphical User Interface (GUI) agents offer cross-platform solutions for automating complex digital tasks, with significant potential to transform productivity workflows. However, their performance is often constrained by the scarcity of…

Artificial Intelligence · Computer Science 2025-04-16 Junlei Zhang , Zichen Ding , Chang Ma , Zijie Chen , Qiushi Sun , Zhenzhong Lan , Junxian He

Many real-world tasks require an agent to reason jointly over text and visual objects, (e.g., navigating in public spaces), which we refer to as context-sensitive text-rich visual reasoning. Specifically, these tasks require an…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Rohan Wadhawan , Hritik Bansal , Kai-Wei Chang , Nanyun Peng

Vision-Language-Action (VLA) models often fail to generalize to unseen camera viewpoints, a limitation stemming from their difficulty in inferring robust 3D geometry from 2D images. We introduce GeoAware-VLA, a simple yet effective approach…

Robotics · Computer Science 2026-03-10 Ali Abouzeid , Malak Mansour , Qinbo Sun , Zezhou Sun , Dezhen Song

Geometric Problem Solving (GPS) poses a unique challenge for Multimodal Large Language Models (MLLMs), requiring not only the joint interpretation of text and diagrams but also iterative visuospatial reasoning. While existing approaches…

Artificial Intelligence · Computer Science 2026-03-26 Shichao Weng , Zhiqiang Wang , Yuhua Zhou , Rui Lu , Ting Liu , Zhiyang Teng , Xiaozhang Liu , Hanmeng Liu

Earth Observation (EO) is moving beyond static prediction toward multi-step analytical workflows that require coordinated reasoning over data, tools, and geospatial state. While foundation models and vision-language models have advanced…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Muhammad Akhtar Munir , Muhammad Umer Sheikh , Akashah Shabbir , Muhammad Haris Khan , Fahad Khan , Xiao Xiang Zhu , Begum Demir , Salman Khan

We introduce VisTA, a new reinforcement learning framework that empowers visual agents to dynamically explore, select, and combine tools from a diverse library based on empirical performance. Existing methods for tool-augmented reasoning…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Zeyi Huang , Yuyang Ji , Anirudh Sundara Rajan , Zefan Cai , Wen Xiao , Haohan Wang , Junjie Hu , Yong Jae Lee

This paper introduces GeoChain, a large-scale benchmark for evaluating step-by-step geographic reasoning in multimodal large language models (MLLMs). Leveraging 1.46 million Mapillary street-level images, GeoChain pairs each image with a…

Artificial Intelligence · Computer Science 2025-09-10 Sahiti Yerramilli , Nilay Pande , Rynaa Grover , Jayant Sravan Tamarapalli

Advanced chart question answering requires both precise perception of small visual elements and multi-step reasoning across several subplots. While existing MLLMs are strong at understanding single plots, they often struggle with multi-step…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Qihua Dong , Ruozhen He , Junwen Chen , Yizhou Wang , Xu Ma , Songyao Jiang , Yun Fu

Accurate localization in diverse environments is a fundamental challenge in computer vision and robotics. The task involves determining a sensor's precise position and orientation, typically a camera, within a given space. Traditional…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Luca Di Giammarino , Boyang Sun , Giorgio Grisetti , Marc Pollefeys , Hermann Blum , Daniel Barath

Autonomous agents capable of planning, reasoning, and executing actions on the web offer a promising avenue for automating computer tasks. However, the majority of existing benchmarks primarily focus on text-based agents, neglecting many…

Multimodal Large Language Models (MLLMs) are increasingly applied in real-world scenarios where user-provided images are often imperfect, requiring active image manipulations such as cropping, editing, or enhancement to uncover salient…

Recently, large language models (LLMs) have demonstrated remarkable problem-solving capabilities by autonomously integrating with external tools for collaborative reasoning. However, due to the inherently complex and diverse nature of…

Artificial Intelligence · Computer Science 2025-11-03 Mengjie Deng , Guanting Dong , Zhicheng Dou

While agent evaluation has shifted toward long-horizon tasks, most benchmarks still emphasize local, step-level reasoning rather than the global constrained optimization (e.g., time and financial budgets) that demands genuine planning…

Artificial Intelligence · Computer Science 2026-01-27 Yinger Zhang , Shutong Jiang , Renhao Li , Jianhong Tu , Yang Su , Lianghao Deng , Xudong Guo , Chenxu Lv , Junyang Lin

Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the lack of relevant benchmarks. Previous work in visual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yiqing Shen , Chenjia Li , Chenxiao Fan , Mathias Unberath

While text-to-image generation has achieved unprecedented fidelity, the vast majority of existing models function fundamentally as static text-to-pixel decoders. Consequently, they often fail to grasp implicit user intentions. Although…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Jun He , Junyan Ye , Zilong Huang , Dongzhi Jiang , Chenjue Zhang , Leqi Zhu , Renrui Zhang , Xiang Zhang , Weijia Li

Active Geo-localization (AGL) is the task of localizing a goal, represented in various modalities (e.g., aerial images, ground-level images, or text), within a predefined search area. Current methods approach AGL as a goal-reaching…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Li Mi , Manon Bechaz , Zeming Chen , Antoine Bosselut , Devis Tuia
‹ Prev 1 3 4 5 6 7 10 Next ›