English
Related papers

Related papers: Survey of Multimodal Geospatial Foundation Models:…

200 papers

With the significant advancements of Large Language Models (LLMs) in the field of Natural Language Processing (NLP), the development of image-text multimodal models has garnered widespread attention. Current surveys on image-text multimodal…

Computation and Language · Computer Science 2024-06-21 Ruifeng Guo , Jingxuan Wei , Linzhuang Sun , Bihui Yu , Guiyong Chang , Dawei Liu , Sibo Zhang , Zhengbing Yao , Mingjun Xu , Liping Bu

Large unimodal foundation models for vision and language encode rich semantic structures, yet aligning them typically requires computationally intensive multimodal fine-tuning. Such approaches depend on large-scale parameter updates, are…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Abhishek Dalvi , Vasant Honavar

Multimodal Vision Language Models (VLMs) have emerged as a transformative topic at the intersection of computer vision and natural language processing, enabling machines to perceive and reason about the world through both visual and textual…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Zongxia Li , Xiyang Wu , Hongyang Du , Fuxiao Liu , Huy Nghiem , Guangyao Shi

Large foundation models, including large language models (LLMs), vision transformers (ViTs), diffusion, and LLM-based multimodal models, are revolutionizing the entire machine learning lifecycle, from training to deployment. However, the…

Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data. Current methods primarily utilize task-specific models, while recent foundation models for…

Machine Learning · Computer Science 2025-10-16 Dominik J. Mühlematter , Lin Che , Ye Hong , Martin Raubal , Nina Wiedemann

Recent work explores new opportunities at the intersection of vision-language-action models (VLAs) and geometric foundation models (GFMs) for 3D reconstruction, such as VGGT. While the resulting geometric VLAs often show improved…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Yurou Yang , Muyuan Lin , Roberto Martin-Martin , Martin Labrie , Shreekant Gayaka , Cheng-Hao Kuo , Luca Carlone

Precisely perceiving the geometric and semantic properties of real-world 3D objects is crucial for the continued evolution of augmented reality and robotic applications. To this end, we present Foundation Model Embedded Gaussian Splatting…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Xingxing Zuo , Pouya Samangouei , Yunwen Zhou , Yan Di , Mingyang Li

Autoformalization, which translates natural language mathematics into formal statements to enable machine reasoning, faces fundamental challenges in the wild due to the multimodal nature of the physical world, where physics requires…

Foundation models or pre-trained models have substantially improved the performance of various language, vision, and vision-language understanding tasks. However, existing foundation models can only perform the best in one type of tasks,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Xinsong Zhang , Yan Zeng , Jipeng Zhang , Hang Li

Foundation models pretrained on diverse data at scale have demonstrated extraordinary capabilities in a wide range of vision and language tasks. When such models are deployed in real world environments, they inevitably interface with other…

Artificial Intelligence · Computer Science 2023-03-08 Sherry Yang , Ofir Nachum , Yilun Du , Jason Wei , Pieter Abbeel , Dale Schuurmans

The mainstream paradigm of remote sensing image interpretation has long been dominated by vision-centered models, which rely on visual features for semantic understanding. However, these models face inherent limitations in handling…

Artificial Intelligence · Computer Science 2026-01-28 Haifeng Li , Wang Guo , Haiyang Wu , Mengwei Wu , Jipeng Zhang , Qing Zhu , Yu Liu , Xin Huang , Chao Tao

Geospatial technologies are becoming increasingly essential in our world for a wide range of applications, including agriculture, urban planning, and disaster response. To help improve the applicability and performance of deep learning…

Computer Vision and Pattern Recognition · Computer Science 2023-09-04 Matias Mendieta , Boran Han , Xingjian Shi , Yi Zhu , Chen Chen

Multimodal Generative Models (MGMs) have rapidly evolved beyond text generation, now spanning diverse output modalities including images, music, video, human motion, and 3D objects, by integrating language with other sensory modalities…

Multimedia · Computer Science 2025-11-25 Longzhen Han , Awes Mubarak , Almas Baimagambetov , Nikolaos Polatidis , Thar Baker

Vision systems to see and reason about the compositional nature of visual scenes are fundamental to understanding our world. The complex relations between objects and their locations, ambiguities, and variations in the real-world…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Muhammad Awais , Muzammal Naseer , Salman Khan , Rao Muhammad Anwer , Hisham Cholakkal , Mubarak Shah , Ming-Hsuan Yang , Fahad Shahbaz Khan

Recent developments in multimodal methodologies have marked the beginning of an exciting era for models adept at processing diverse data types, encompassing text, audio, and visual content. Models like GPT-4V, which merge computer vision…

Computation and Language · Computer Science 2024-11-15 Xiang Zhang , Senyu Li , Ning Shi , Bradley Hauer , Zijun Wu , Grzegorz Kondrak , Muhammad Abdul-Mageed , Laks V. S. Lakshmanan

Foundation models leverage large-scale pretraining to capture extensive knowledge, demonstrating generalization in a wide range of language tasks. By comparison, vision foundation models (VFMs) often exhibit uneven improvements across…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Shiqi Huang , Yipei Wang , Natasha Thorley , Alexander Ng , Shaheer Saeed , Mark Emberton , Shonit Punwani , Veeru Kasivisvanathan , Dean Barratt , Daniel Alexander , Yipeng Hu

Systems that can find correspondences between multiple modalities, such as between speech and images, have great potential to solve different recognition and data analysis tasks in an unsupervised manner. This work studies multimodal…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Khazar Khorrami , Okko Räsänen

Recent progress in self-supervision has shown that pre-training large neural networks on vast amounts of unsupervised data can lead to substantial increases in generalization to downstream tasks. Such models, recently coined foundation…

Artificial intelligence is a key enabler for next-generation wireless communication and sensing. Yet, today's learning-based wireless techniques do not generalize well: most models are task-specific, environment-dependent, and limited to…

Signal Processing · Electrical Eng. & Systems 2026-02-05 Vahid Yazdnian , Yasaman Ghasempour

Vision foundation models (VFMs) are predominantly developed using data-centric methods. These methods require training on vast amounts of data usually with high-quality labels, which poses a bottleneck for most institutions that lack both…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Jiabo Huang , Chen Chen , Lingjuan Lyu
‹ Prev 1 4 5 6 7 8 10 Next ›