English
Related papers

Related papers: Survey of Multimodal Geospatial Foundation Models:…

200 papers

Foundation models (FMs) are large-scale deep learning models trained on massive datasets, often using self-supervised learning techniques. These models serve as a versatile base for a wide range of downstream tasks, including those in…

Machine Learning · Computer Science 2025-01-17 Wasif Khan , Seowung Leem , Kyle B. See , Joshua K. Wong , Shaoting Zhang , Ruogu Fang

The prevalence of Vision-Language Models (VLMs) raises important questions about privacy in an era where visual information is increasingly available. While foundation VLMs demonstrate broad knowledge and learned capabilities, we…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Neel Jay , Hieu Minh Nguyen , Trung Dung Hoang , Jacob Haimes

Remote sensing has evolved from simple image acquisition to complex systems capable of integrating and processing visual and textual data. This review examines the development and application of multi-modal language models (MLLMs) in remote…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Xintian Sun , Benji Peng , Charles Zhang , Fei Jin , Qian Niu , Junyu Liu , Keyu Chen , Ming Li , Pohsun Feng , Ziqian Bi , Ming Liu , Xinyuan Song , Yichao Zhang

Given the ubiquity of graph data and its applications in diverse domains, building a Graph Foundation Model (GFM) that can work well across different graphs and tasks with a unified backbone has recently garnered significant interests. A…

Machine Learning · Computer Science 2024-06-18 Zhikai Chen , Haitao Mao , Jingzhe Liu , Yu Song , Bingheng Li , Wei Jin , Bahare Fatemi , Anton Tsitsulin , Bryan Perozzi , Hui Liu , Jiliang Tang

Visual grounding refers to the ability of a model to identify a region within some visual input that matches a textual description. Consequently, a model equipped with visual grounding capabilities can target a wide range of applications in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Georgios Pantazopoulos , Eda B. Özyiğit

This manuscript explores multimodal alignment, translation, fusion, and transference to enhance machine understanding of complex inputs. We organize the work into five chapters, each addressing unique challenges in multimodal machine…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Gorjan Radevski

Foundation models (FMs) have emerged as a powerful paradigm, enabling a diverse range of data analytics and knowledge discovery tasks across scientific fields. Inspired by the success of FMs, particularly large language models, researchers…

Machine Learning · Computer Science 2025-11-27 Sean Bin Yang , Ying Sun , Yunyao Cheng , Yan Lin , Kristian Torp , Jilin Hu

The advent of foundation models, which are pre-trained on vast datasets, has ushered in a new era of computer vision, characterized by their robustness and remarkable zero-shot generalization capabilities. Mirroring the transformative…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Xu Liu , Tong Zhou , Yuanxin Wang , Yuping Wang , Qinjingwen Cao , Weizhi Du , Yonghuan Yang , Junjun He , Yu Qiao , Yiqing Shen

Large-scale pretraining on Earth observation imagery has yielded powerful representations of the natural and built environment. However, most existing geospatial foundation models do not directly model the structured socioeconomic…

Machine Learning · Computer Science 2026-05-15 Yuhao Liu , Sadeer Al-Kindi , Ashok Veeraraghavan , Guha Balakrishnan

Foundation models like ChatGPT and Sora that are trained on a huge scale of data have made a revolutionary social impact. However, it is extremely challenging for sensors in many different fields to collect similar scales of natural images…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Chenyang Lei , Liyi Chen , Jun Cen , Xiao Chen , Zhen Lei , Felix Heide , Qifeng Chen , Zhaoxiang Zhang

Geospatial foundation models (GFMs) have been proposed as generalizable backbones for disaster response, land-cover mapping, food-security monitoring, and other high-stakes Earth-observation tasks. Yet the published work about these models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Isaac Corley , Nils Lehmann , Caleb Robinson , Gabriel Tseng , Anthony Fuller , Hamed Alemohammad , Evan Shelhamer , Jennifer Marcus , Hannah Kerner

The advent of foundation models has revolutionized the fields of natural language processing and computer vision, paving the way for their application in autonomous driving (AD). This survey presents a comprehensive review of more than 40…

Machine Learning · Computer Science 2024-09-06 Haoxiang Gao , Zhongruo Wang , Yaqian Li , Kaiwen Long , Ming Yang , Yiqing Shen

This survey and application guide to multimodal large language models(MLLMs) explores the rapidly developing field of MLLMs, examining their architectures, applications, and impact on AI and Generative Models. Starting with foundational…

Artificial Intelligence · Computer Science 2025-12-02 Chia Xin Liang , Pu Tian , Caitlyn Heqi Yin , Yao Yua , Wei An-Hou , Li Ming , Xinyuan Song , Tianyang Wang , Ziqian Bi , Ming Liu

Multimodal foundation models aim to create a unified representation space that abstracts away from surface features like language syntax or modality differences. To investigate this, we study the internal representations of three recent…

Computation and Language · Computer Science 2025-02-21 Hyunji Lee , Danni Liu , Supriti Sinhamahapatra , Jan Niehues

Foundation models have emerged as a powerful paradigm in computational pathology (CPath), enabling scalable and generalizable analysis of histopathological images. While early developments centered on uni-modal models trained solely on…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Dong Li , Guihong Wan , Xintao Wu , Xinyu Wu , Xiaohui Chen , Yi He , Christine G. Lian , Peter K. Sorger , Yevgeniy R. Semenov , Chen Zhao

The rapid development of Vision Foundation Models (VFMs), particularly Vision Transformers (ViT) and Segment Anything Model (SAM), has sparked significant advances in the field of medical image analysis. These models have demonstrated…

Image and Video Processing · Electrical Eng. & Systems 2025-02-24 Pengchen Liang , Bin Pu , Haishan Huang , Yiwei Li , Hualiang Wang , Weibo Ma , Qing Chang

Cross-modal artificial intelligence, represented by visual language models, has achieved significant success in general image understanding. However, a fundamental cognitive inconsistency exists between general visual representation and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Yi Yang , Xiaokun Zhang , Qingchen Fang , Jing Liu , Ziqi Ye , Rui Li , Li Liu , Haipeng Wang

The emergence of Large Language Models (LLMs) and multimodal foundation models (FMs) has generated heightened interest in their applications that integrate vision and language. This paper investigates the capabilities of ChatGPT-4V and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Zhenyuan Yang , Xuhui Lin , Qinyi He , Ziye Huang , Zhengliang Liu , Hanqi Jiang , Peng Shu , Zihao Wu , Yiwei Li , Stephen Law , Gengchen Mai , Tianming Liu , Tao Yang

Remote sensing (RS) techniques are increasingly crucial for deepening our understanding of the planet. As the volume and diversity of RS data continue to grow exponentially, there is an urgent need for advanced data modeling and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Danfeng Hong , Chenyu Li , Xuyang Li , Gustau Camps-Valls , Jocelyn Chanussot

In recent years, Graph Foundation Models (GFMs) have gained significant attention for their potential to generalize across diverse graph domains and tasks. Some works focus on Domain-Specific GFMs, which are designed to address a variety of…

Machine Learning · Computer Science 2025-03-13 Yuxiang Wang , Wenqi Fan , Suhang Wang , Yao Ma
‹ Prev 1 3 4 5 6 7 10 Next ›