English
Related papers

Related papers: Falcon: A Remote Sensing Vision-Language Foundatio…

200 papers

Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human vision. Computer vision foundation models, which are trained on…

Precise delineation of anatomical and pathological structures within 3D medical volumes is crucial for accurate diagnosis, effective surgical planning, and longitudinal disease monitoring. Despite advancements in AI, clinically viable…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Abdur R. Fayjie , Pankhi Kashyap , Jutika Borah , Patrick Vandewalle

Remote sensing imagery presents vast, inherently unstructured spatial data, necessitating sophisticated reasoning to interpret complex user intents and contextual relationships beyond simple recognition tasks. In this paper, we aim to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Liang Yao , Fan Liu , Hongbo Lu , Chuanyi Zhang , Rui Min , Shengxiang Xu , Shimin Di , Pai Peng

Quantitative remote sensing inversion plays a critical role in environmental monitoring, enabling the estimation of key ecological variables such as vegetation indices, canopy structure, and carbon stock. Although vision foundation models…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Zhenyu Yu , Mohd. Yamani Idna Idris , Pei Wang

Executing multiple tasks simultaneously in medical image analysis, including segmentation, classification, detection, and regression, often introduces significant challenges regarding model generalizability and the optimization of shared…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Hui Wan , Libin Lan

Vision-language modeling (VLM) aims to bridge the information gap between images and natural language. Under the new paradigm of first pre-training on massive image-text pairs and then fine-tuning on task-specific data, VLM in the remote…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Xingxing Weng , Chao Pang , Gui-Song Xia

Advanced interpretation of hyperspectral remote sensing images benefits many precise Earth observation tasks. Recently, visual foundation models have promoted the remote sensing interpretation but concentrating on RGB and multispectral…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Jingtao Li , Yingyi Liu , Xinyu Wang , Yunning Peng , Chen Sun , Shaoyu Wang , Zhendong Sun , Tian Ke , Xiao Jiang , Tangwei Lu , Anran Zhao , Yanfei Zhong

Remote Sensing Vision-Language Models (RS VLMs) have made much progress in the tasks of remote sensing (RS) image comprehension. While performing well in multi-modal reasoning and multi-turn conversations, the existing models lack…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Xu Liu , Zhouhui Lian

Remote Sensing Image Captioning (RSIC) is a cross-modal field bridging vision and language, aimed at automatically generating natural language descriptions of features and scenes in remote sensing imagery. Despite significant advances in…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Qing Zhou , Tao Yang , Junyu Gao , Weiping Ni , Junzheng Wu , Qi Wang

Remote sensing has become a vital tool across sectors such as urban planning, environmental monitoring, and disaster response. While the volume of data generated has increased significantly, traditional vision models are often constrained…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Jia Yun Chua , Argyrios Zolotas , Miguel Arana-Catania

This work introduces Falcon-H1R, a 7B-parameter reasoning-optimized model that establishes the feasibility of achieving competitive reasoning performance with small language models (SLMs). Falcon-H1R stands out for its parameter efficiency,…

Multimodal large language models (MLLMs) demonstrate strong perception and reasoning performance on existing remote sensing (RS) benchmarks. However, most prior benchmarks rely on low-resolution imagery, and some high-resolution benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Yunkai Dang , Meiyi Zhu , Donghao Wang , Yizhuo Zhang , Jiacheng Yang , Qi Fan , Yuekun Yang , Wenbin Li , Feng Miao , Yang Gao

As a powerful all-weather Earth observation tool, synthetic aperture radar (SAR) remote sensing enables critical military reconnaissance, maritime surveillance, and infrastructure monitoring. Although Vision language models (VLMs) have made…

Computation and Language · Computer Science 2025-03-05 Zhiming Ma , Xiayang Xiao , Sihao Dong , Peidong Wang , HaiPeng Wang , Qingyun Pan

We present FoundAtion-model-guided decoupled LoCO-maNipulation visuomotor policies (FALCON), a framework for loco-manipulation that combines modular diffusion policies with a vision-language foundation model as the coordinator. Our approach…

Robotics · Computer Science 2025-12-05 Chengyang He , Ge Sun , Yue Bai , Junkai Lu , Jiadong Zhao , Guillaume Sartoretti

Machine-learning algorithms have shown outstanding image recognition or classification performance for computer vision applications. However, the compute and energy requirement for implementing such classifier models for large-scale…

Computer Vision and Pattern Recognition · Computer Science 2017-03-09 Priyadarshini Panda , Aayush Ankit , Parami Wijesinghe , Kaushik Roy

Today's unsupervised image segmentation algorithms often segment suboptimally. Modern graph-cut based approaches rely on high-dimensional attention maps from Transformer-based foundation models, typically employing a relaxed Normalized Cut…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Xiao Zhang , Xiangyu Han , Xiwen Lai , Yao Sun , Pei Zhang , Konrad Kording

To navigate safely and efficiently in crowded spaces, robots should not only perceive the current state of the environment but also anticipate future human movements. In this paper, we propose a reinforcement learning architecture, namely…

Robotics · Computer Science 2025-02-11 Zeying Gong , Tianshuai Hu , Ronghe Qiu , Junwei Liang

Foundation models are becoming increasingly effective in the medical domain, offering pre-trained models on large datasets that can be readily adapted for downstream tasks. Despite progress, fetal ultrasound images remain a challenging…

Adapting vision-language models to remote sensing imagery remains challenging due to two key factors: limited semantic coverage in textual representations and insufficient adaptability of visual features. These issues are particularly…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Yu Hu , Jianyang Gu , Hao Liu , Yue Cao , Jozsef Hamari , Zheng Liu , Mohsen Zardadi

Remote Sensing (RS) is a crucial technology for observing, monitoring, and interpreting our planet, with broad applications across geoscience, economics, humanitarian fields, etc. While artificial intelligence (AI), particularly deep…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Aoran Xiao , Weihao Xuan , Junjue Wang , Jiaxing Huang , Dacheng Tao , Shijian Lu , Naoto Yokoya