English
Related papers

Related papers: Multimodal Wireless Foundation Models

200 papers

Humanoid robots are drawing significant attention as versatile platforms for complex motor control, human-robot interaction, and general-purpose physical intelligence. However, achieving efficient whole-body control (WBC) in humanoids…

Robotics · Computer Science 2026-02-10 Mingqi Yuan , Tao Yu , Wenqi Ge , Xiuyong Yao , Huijiang Wang , Jiayu Chen , Bo Li , Wei Zhang , Wenjun Zeng , Hua Chen , Xin Jin

We present a novel approach in the domain of federated learning (FL), particularly focusing on addressing the challenges posed by modality heterogeneity, variability in modality availability across clients, and the prevalent issue of…

Machine Learning · Computer Science 2023-12-19 Minh Tran , Roochi Shah , Zejun Gong

The Web of Things (WoT) enhances interoperability across web-based and ubiquitous computing platforms while complementing existing IoT standards. The multimodal Federated Learning (FL) paradigm has been introduced to enhance WoT by enabling…

Machine Learning · Computer Science 2025-02-25 Yi Liu , Cong Wang , Xingliang Yuan

Leading approaches in machine vision employ different architectures for different tasks, trained on costly task-specific labeled datasets. This complexity has held back progress in areas, such as robotics, where robust task-general…

Computer Vision and Pattern Recognition · Computer Science 2023-06-06 Daniel M. Bear , Kevin Feigelis , Honglin Chen , Wanhee Lee , Rahul Venkatesh , Klemen Kotar , Alex Durango , Daniel L. K. Yamins

Federated learning (FL) enables the collaborative training of deep neural networks across decentralized data archives (i.e., clients) without sharing the local data of the clients. Most of the existing FL methods assume that the data…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Barış Büyüktaş , Gencer Sumbul , Begüm Demir

Foundation models have garnered increasing attention for representation learning in remote sensing. Many such foundation models adopt approaches that have demonstrated success in computer vision with minimal domain-specific modification.…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Kevin Lane , Morteza Karimzadeh

We present Masked Frequency Modeling (MFM), a unified frequency-domain-based approach for self-supervised pre-training of visual models. Instead of randomly inserting mask tokens to the input embeddings in the spatial domain, in this paper,…

Computer Vision and Pattern Recognition · Computer Science 2023-04-26 Jiahao Xie , Wei Li , Xiaohang Zhan , Ziwei Liu , Yew Soon Ong , Chen Change Loy

Multimodal deep learning systems which employ multiple modalities like text, image, audio, video, etc., are showing better performance in comparison with individual modalities (i.e., unimodal) systems. Multimodal machine learning involves…

Machine Learning · Computer Science 2022-01-19 Anil Rahate , Rahee Walambe , Sheela Ramanna , Ketan Kotecha

We present VisionFM, a foundation model pre-trained with 3.4 million ophthalmic images from 560,457 individuals, covering a broad range of ophthalmic diseases, modalities, imaging devices, and demography. After pre-training, VisionFM…

Automatic modulation classification (AMC) in real-world deployments demands robustness to distribution shifts arising from hardware impairments, unseen propagation environments, and recording conditions never encountered during training.…

Machine Learning · Computer Science 2026-05-06 Md Raihan Uddin , Tolunay Seyfi , Fatemeh Afghah

Multimodal learning aims to build models that can process and relate information from multiple modalities. Despite years of development in this field, it still remains challenging to design a unified network for processing various…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Yiyuan Zhang , Kaixiong Gong , Kaipeng Zhang , Hongsheng Li , Yu Qiao , Wanli Ouyang , Xiangyu Yue

Beamforming (BF) is essential for enhancing system capacity in fifth generation (5G) and beyond wireless networks, yet exhaustive beam training in ultra-massive multiple-input multiple-output (MIMO) systems incurs substantial overhead. To…

Signal Processing · Electrical Eng. & Systems 2026-02-11 Yanliang Jin , Yunfan Li , Jiang Jun , Yuan Gao , Shengli Liu , Jianbo Du , Zhaohui Yang , Shugong Xu

Recent studies have focused on utilizing multi-modal data to develop robust models for facial Action Unit (AU) detection. However, the heterogeneity of multi-modal data poses challenges in learning effective representations. One such…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Xiang Zhang , Huiyuan Yang , Taoyue Wang , Xiaotian Li , Lijun Yin

Integrated sensing and communication (ISAC) is a key enabler for low-altitude wireless networks (LAWNs), providing simultaneous environmental perception and data transmission in complex aerial scenarios. By combining heterogeneous sensing…

Signal Processing · Electrical Eng. & Systems 2025-12-02 Kai Zhang , Wentao Yu , Hengtao He , Shenghui Song , Jun Zhang , Khaled B. Letaief

Wireless channel foundation model (WCFM) is a task-agnostic AI model that is pre-trained to learn a universal channel representation for a wide range of communications and sensing tasks. While existing works on WCFM have demonstrated its…

Signal Processing · Electrical Eng. & Systems 2026-01-09 Yuwei Wang , Li Sun , Tingting Yang , Yuxuan Shi , Maged Elkashlan , Xiao Tang

Foundation models or pre-trained models have substantially improved the performance of various language, vision, and vision-language understanding tasks. However, existing foundation models can only perform the best in one type of tasks,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Xinsong Zhang , Yan Zeng , Jipeng Zhang , Hang Li

Facial expression recognition is an essential task for various applications, including emotion detection, mental health analysis, and human-machine interactions. In this paper, we propose a multi-modal facial expression recognition method…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Jun-Hwa Kim , Namho Kim , Chee Sun Won

Large Language Models (LLMs) have made the ambitious quest for generalist agents significantly far from being a fantasy. A key hurdle for building such general models is the diversity and heterogeneity of tasks and modalities. A promising…

Computer Vision and Pattern Recognition · Computer Science 2023-12-25 Mustafa Shukor , Corentin Dancette , Alexandre Rame , Matthieu Cord

Multimodal learning enhances the perceptual capabilities of cognitive systems by integrating information from different sensory modalities. However, existing multimodal fusion research typically assumes static integration, not fully…

Neural and Evolutionary Computing · Computer Science 2025-05-16 Xiang He , Dongcheng Zhao , Yang Li , Qingqun Kong , Xin Yang , Yi Zeng

Multimodal semantic communication has gained widespread attention due to its ability to enhance downstream task performance. A key challenge in such systems is the effective fusion of features from different modalities, which requires the…

Image and Video Processing · Electrical Eng. & Systems 2025-09-03 Haoshuo Zhang , Yufei Bo , Hongwei Zhang , Meixia Tao