English
Related papers

Related papers: Geospatial-Temporal Sensemaking of Remote Sensing …

200 papers

Text-Centric Visual Question Answering (TEC-VQA) in its proper format not only facilitates human-machine interaction in text-centric visual environments but also serves as a de facto gold proxy to evaluate AI models in the domain of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Jingqun Tang , Qi Liu , Yongjie Ye , Jinghui Lu , Shu Wei , Chunhui Lin , Wanqing Li , Mohamad Fitri Faiz Bin Mahmood , Hao Feng , Zhen Zhao , Yangfan He , Kuan Lu , Yanjie Wang , Yuliang Liu , Hao Liu , Xiang Bai , Can Huang

The entorhinal-hippocampal formation is the mammalian brain's navigation system, encoding both physical and abstract spaces via grid cells. This system is well-studied in neuroscience, and its efficiency and versatility make it attractive…

Neural and Evolutionary Computing · Computer Science 2025-03-12 Sven Krausse , Emre Neftci , Friedrich T. Sommer , Alpha Renner

This paper tackles the high computational/space complexity associated with Multi-Head Self-Attention (MHSA) in vanilla vision transformers. To this end, we propose Hierarchical MHSA (H-MHSA), a novel approach that computes self-attention in…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Yun Liu , Yu-Huan Wu , Guolei Sun , Le Zhang , Ajad Chhatkuli , Luc Van Gool

This paper develops a Versatile and Honest vision language Model (VHM) for remote sensing image analysis. VHM is built on a large-scale remote sensing image-text dataset with rich-content captions (VersaD), and an honest instruction dataset…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Chao Pang , Xingxing Weng , Jiang Wu , Jiayu Li , Yi Liu , Jiaxing Sun , Weijia Li , Shuai Wang , Litong Feng , Gui-Song Xia , Conghui He

Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate…

Visual Question Answering (VQA) is a challenging task that requires systems to provide accurate answers to questions based on image content. Current VQA models struggle with complex questions due to limitations in capturing and integrating…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Peiyuan Chen , Zecheng Zhang , Yiping Dong , Li Zhou , Han Wang

Document Visual Question Answering (VQA) aims to understand visually-rich documents to answer questions in natural language, which is an emerging research topic for both Natural Language Processing and Computer Vision. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2023-05-05 Fengbin Zhu , Wenqiang Lei , Fuli Feng , Chao Wang , Haozhou Zhang , Tat-Seng Chua

Driven by practical demands in land resource monitoring and national defense security, this paper introduces the Remote Sensing Copy-Move Question Answering (RSCMQA) task. Unlike traditional Remote Sensing Visual Question Answering (RSVQA),…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Ze Zhang , Enyuan Zhao , Di Niu , Jie Nie , Xinyue Liang , Lei Huang

Remote sensing visual question answering (RSVQA) opens new opportunities for the use of overhead imagery by the general public, by enabling human-machine interaction with natural language. Building on the recent advances in natural language…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Christel Chappuis , Eliot Walt , Vincent Mendez , Sylvain Lobry , Bertrand Le Saux , Devis Tuia

Deep learning approaches have shown promising results in remote sensing high spatial resolution (HSR) land-cover mapping. However, urban and rural scenes can show completely different geographical landscapes, and the inadequate…

Computer Vision and Pattern Recognition · Computer Science 2022-06-01 Junjue Wang , Zhuo Zheng , Ailong Ma , Xiaoyan Lu , Yanfei Zhong

Although large-scale video-language pre-training models, which usually build a global alignment between the video and the text, have achieved remarkable progress on various downstream tasks, the idea of adopting fine-grained information…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Weihong Zhong , Mao Zheng , Duyu Tang , Xuan Luo , Heng Gong , Xiaocheng Feng , Bing Qin

Vision-language-action (VLA) models have shown strong semantic grounding and task generalization in manipulation, but aerial deployment remains difficult because drones require low-latency closed-loop guidance under strict onboard compute…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Justin williams , Kishor Datta Gupta , Roy George , Mrinmoy Sarkar

Three-dimensional geospatial analysis is critical for applications in urban planning, climate adaptation, and environmental assessment. However, current methodologies depend on costly, specialized sensors, such as LiDAR and multispectral…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Mai Tsujimoto , Junjue Wang , Weihao Xuan , Naoto Yokoya

The Audio Visual Question Answering (AVQA) task aims to answer questions related to various visual objects, sounds, and their interactions in videos. Such naturally multimodal videos contain rich and complex dynamic audio-visual components,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Guangyao Li , Henghui Du , Di Hu

Traditional deep learning methods struggle to simultaneously segment, recognize, and forecast human activities from sensor data. This limits their usefulness in many fields such as healthcare and assisted living, where real-time…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Shuangjian Li , Tao Zhu , Mingxing Nie , Huansheng Ning , Zhenyu Liu , Liming Chen

This paper proposes deep convolutional network models that utilize local and global context to make human activity label predictions in still images, achieving state-of-the-art performance on two recent datasets with hundreds of labels…

Computer Vision and Pattern Recognition · Computer Science 2016-07-29 Arun Mallya , Svetlana Lazebnik

Efficient long-short temporal modeling is key for enhancing the performance of action recognition task. In this paper, we propose a new two-stream action recognition network, termed as MENet, consisting of a Motion Enhancement (ME) module…

Computer Vision and Pattern Recognition · Computer Science 2021-07-01 Liyu Wu , Yuexian Zou , Can Zhang

Given the remarkable success that large visual language models (LVLMs) have achieved in image perception tasks, the endeavor to make LVLMs perceive the world like humans is drawing increasing attention. Current multi-modal benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Siwei Wu , Kang Zhu , Yu Bai , Yiming Liang , Yizhi Li , Haoning Wu , J. H. Liu , Ruibo Liu , Xingwei Qu , Xuxin Cheng , Ge Zhang , Wenhao Huang , Chenghua Lin

We present a novel method, AutoSpatial, an efficient approach with structured spatial grounding to enhance VLMs' spatial reasoning. By combining minimal manual supervision with large-scale Visual Question-Answering (VQA) pairs…

Robotics · Computer Science 2026-05-05 Yangzhe Kong , Daeun Song , Jing Liang , Dinesh Manocha , Ziyu Yao , Xuesu Xiao