English

M3F-UAV: A Missing-Modality Multimodal Foundation Model for Low-Altitude Wireless Sensing

Signal Processing 2026-07-15 v1

Abstract

Low-altitude unmanned aerial vehicles (UAVs) are emerging as key platforms for wireless intelligence tasks. However, practical low-altitude wireless systems usually operate in complex urban environments, where visual occlusion, sparse geometric observations, multipath propagation, and sensor failures may degrade the reliability of single-modality models. To address these challenges, this paper proposes M3F-UAV, a missing-modality multimodal foundation model for low-altitude wireless sensing. The proposed framework learns a unified multimodal representation from visual, geometric, and wireless observations. Specifically, modality-specific pretrained feature extractors are adopted for RGB/depth images, LiDAR point clouds, and CSI matrices, respectively. Through cross-modal fusion and missing-modality-aware pretraining with feature-level masked reconstruction and UAV localization objectives, M3F-UAV can extract fixed-size features from different modality combinations and adapt them to downstream low-altitude wireless tasks with lightweight task heads. Experiments on the LAMBDA dataset show that M3F-UAV outperforms single-modality baselines and maintains robust performance under missing-modality settings.

Cite

@article{arxiv.2607.13678,
  title  = {M3F-UAV: A Missing-Modality Multimodal Foundation Model for Low-Altitude Wireless Sensing},
  author = {Pengxuan Gao and Kai Ying and Botao Wu and Jianhua Mo and Qingsong Wen},
  journal= {arXiv preprint arXiv:2607.13678},
  year   = {2026}
}