English
Related papers

Related papers: TerraMind: Large-Scale Generative Multimodality fo…

200 papers

Humanoid robot technology is advancing rapidly, with manufacturers introducing diverse heterogeneous visual perception modules tailored to specific scenarios. Among various perception paradigms, occupancy-based representation has become…

Earth observation (EO) spans a broad spectrum of modalities, including optical, radar, multispectral, and hyperspectral data, each capturing distinct environmental signals. However, current vision-language models in EO, particularly…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Zhitong Xiong , Yi Wang , Weikang Yu , Adam J Stewart , Jie Zhao , Nils Lehmann , Thomas Dujardin , Zhenghang Yuan , Pedram Ghamisi , Xiao Xiang Zhu

Earth observation applications increasingly rely on data from multiple sensors, including optical, radar, elevation, and land-cover. Relationships between modalities are fundamental for data integration but are inherently non-injective:…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Miguel Espinosa , Eva Gmelich Meijling , Valerio Marsocci , Elliot J. Crowley , Mikolaj Czerkawski

We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks and low-resolution…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Shufan Li , Jiuxiang Gu , Kangning Liu , Zhe Lin , Zijun Wei , Aditya Grover , Jason Kuen

Multi-modal fusion of sensors is a commonly used approach to enhance the performance of odometry estimation, which is also a fundamental module for mobile robots. However, the question of \textit{how to perform fusion among different…

Robotics · Computer Science 2025-03-20 Leyuan Sun , Guanqun Ding , Yue Qiu , Yusuke Yoshiyasu , Fumio Kanehiro

Large-scale pretraining on Earth observation imagery has yielded powerful representations of the natural and built environment. However, most existing geospatial foundation models do not directly model the structured socioeconomic…

Machine Learning · Computer Science 2026-05-15 Yuhao Liu , Sadeer Al-Kindi , Ashok Veeraraghavan , Guha Balakrishnan

This study aims to address the problem of incomplete information in unimodal images for semantic segmentation and object detection tasks. Existing multimodal fusion methods suffer from limited capability in discriminative modeling of…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Yuchan Jie , Yushen Xu , Xiaosong Li , Huafeng Li , Haishu Tan , Feiping Nie

The human ability to seamlessly perform multimodal reasoning and physical interaction in the open world is a core goal for general purpose embodied intelligent systems. Recent vision-language-action (VLA) models, which are co-trained on…

Multimodal representation learning has shown promising improvements on various vision-language tasks. Most existing methods excel at building global-level alignment between vision and language while lacking effective fine-grained image-text…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Zijia Zhao , Longteng Guo , Xingjian He , Shuai Shao , Zehuan Yuan , Jing Liu

Dense pixel-specific representation learning at scale has been bottlenecked due to the unavailability of large-scale multi-view datasets. Current methods for building effective pretraining datasets heavily rely on annotated 3D meshes, point…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Kalyani Marathe , Mahtab Bigverdi , Nishat Khan , Tuhin Kundu , Patrick Howe , Sharan Ranjit S , Anand Bhattad , Aniruddha Kembhavi , Linda G. Shapiro , Ranjay Krishna

3D content inherently encompasses multi-modal characteristics and can be projected into different modalities (e.g., RGB images, RGBD, and point clouds). Each modality exhibits distinct advantages in 3D asset modeling: RGB images contain…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Ziang Cao , Zhaoxi Chen , Liang Pan , Ziwei Liu

With the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress. However, limited by existing datasets,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Ruiyuan Lyu , Jingli Lin , Tai Wang , Shuai Yang , Xiaohan Mao , Yilun Chen , Runsen Xu , Haifeng Huang , Chenming Zhu , Dahua Lin , Jiangmiao Pang

Large pre-trained models, also known as foundation models (FMs), are trained in a task-agnostic manner on large-scale data and can be adapted to a wide range of downstream tasks by fine-tuning, few-shot, or even zero-shot learning. Despite…

Artificial Intelligence · Computer Science 2023-04-17 Gengchen Mai , Weiming Huang , Jin Sun , Suhang Song , Deepak Mishra , Ninghao Liu , Song Gao , Tianming Liu , Gao Cong , Yingjie Hu , Chris Cundy , Ziyuan Li , Rui Zhu , Ni Lao

How to enable agents to predict the outcomes of their own motion intentions in three-dimensional space has been a fundamental problem in embodied intelligence. To explore general spatial imagination capability, we present AirScape, the…

Masked Image Modeling (MIM) has become an essential method for building foundational visual models in remote sensing (RS). However, the limitations in size and diversity of existing RS datasets restrict the ability of MIM methods to learn…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Fengxiang Wang , Hongzhen Wang , Di Wang , Zonghao Guo , Zhenyu Zhong , Long Lan , Wenjing Yang , Jing Zhang

Cross-modal image-to-image translation among Electro-Optical (EO), Infrared (IR), and Synthetic Aperture Radar (SAR) sensors is essential for comprehensive multi-modal aerial-view analysis. However, translating between these modalities is…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Zhenyuan Chen , Guanyuan Shen , Feng Zhang

Advances in data collection enable the capture of rich patient-generated data: from passive sensing (e.g., wearables and smartphones) to active self-reports (e.g., cross-sectional surveys and ecological momentary assessments). Although…

A multimodal system with Poisson, Gaussian, and multinomial observations is considered. A generative graphical model that combines multiple modalities through common factor loadings is proposed. In this model, latent factors are like…

Applications · Statistics 2015-08-04 Yasin Yilmaz , Alfred O. Hero

Recent Multimodal Large Language Models (MLLMs) have shown high potential for spatial reasoning within 3D scenes. However, they typically rely on computationally expensive 3D representations like point clouds or reconstructed Bird's-Eye…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Shuyao Shi , Kang G. Shin

Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency between these two capabilities. We observe that this issue…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yinyi Luo , Wenwen Wang , Hayes Bai , Marios Savvides , Jindong Wang
‹ Prev 1 8 9 10 Next ›