中文
相关论文

相关论文: X-Fi: A Modality-Invariant Foundation Model for Mu…

200 篇论文

Building multisensory AI systems that learn from multiple sensory inputs such as text, speech, video, real-world sensors, wearable devices, and medical data holds great promise for impact in many scientific areas with practical benefits,…

机器学习 · 计算机科学 2024-05-01 Paul Pu Liang

Fusing multi-modal data can improve the performance of deep learning models. However, missing modalities are common for medical data due to patients' specificity, which is detrimental to the performance of multi-modal models in…

图像与视频处理 · 电气工程与系统科学 2023-09-28 Muyu Wang , Shiyu Fan , Yichen Li , Hui Chen

In the health domain, decisions are often based on different data modalities. Thus, when creating prediction models, multimodal fusion approaches that can extract and combine relevant features from different data modalities, can be highly…

人工智能 · 计算机科学 2024-02-20 Mafalda Malafaia , Thalea Schlender , Peter A. N. Bosman , Tanja Alderliesten

Sensor-based Human Activity Recognition (HAR) underpins many ubiquitous and wearable computing applications, yet current models remain limited by scarce labels, sensor heterogeneity, and weak generalization across users, devices, and…

信号处理 · 电气工程与系统科学 2026-04-10 Sizhen Bian , Mengxi Liu , Lala Shakti Swarup Ray , Bo Zhou , Bin Guo , Zhiwen Yu , Thomas Ploetz , Paul Lukowicz , Siyu Yuan , Vitor Fortes Rey

The problem of information fusion from multiple data-sets acquired by multimodal sensors has drawn significant research attention over the years. In this paper, we focus on a particular problem setting consisting of a physical phenomenon or…

机器学习 · 统计学 2018-11-21 Ori Katz , Ronen Talmon , Yu-Lun Lo , Hau-Tieng Wu

Human identification is an important topic in event detection, person tracking, and public security. There have been numerous methods proposed for human identification, such as face identification, person re-identification, and gait…

计算机视觉与模式识别 · 计算机科学 2022-01-03 Zhizhe Liu , Xingxing Zhang , Zhenfeng Zhu , Shuai Zheng , Yao Zhao , Jian Cheng

Information integration from different modalities is an active area of research. Human beings and, in general, biological neural systems are quite adept at using a multitude of signals from different sensory perceptive fields to interact…

神经与进化计算 · 计算机科学 2021-10-05 Shiv Shankar

In recent years, multi-modal fusion has attracted a lot of research interest, both in academia, and in industry. Multimodal fusion entails the combination of information from a set of different types of sensors. Exploiting complementary…

机器学习 · 计算机科学 2020-08-27 Siddharth Roheda , Hamid Krim , Benjamin S. Riggan

Multimodal foundation models have achieved impressive progress across a wide range of vision-language tasks. However, existing approaches often adopt fixed or task-specific fusion strategies, neglecting the intrinsic variability of modality…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Liam Bennett , Mason Clark , Lucas Anderson , Hana Satou , Olivia Martinez

Wireless foundation models (WFMs) have recently demonstrated promising capabilities, jointly performing multiple wireless functions and adapting effectively to new environments. However, while current WFMs process only one modality,…

信号处理 · 电气工程与系统科学 2026-02-20 Ahmed Aboulfotouh , Hatem Abou-Zeid

One of the major reasons for misclassification of multiplex actions during action recognition is the unavailability of complementary features that provide the semantic information about the actions. In different domains these features are…

计算机视觉与模式识别 · 计算机科学 2020-08-25 Zeeshan Ahmad , Naimul Khan

Multimodal Sentiment Analysis is an active area of research that leverages multimodal signals for affective understanding of user-generated videos. The predominant approach, addressing this task, has been to develop sophisticated fusion…

计算与语言 · 计算机科学 2020-10-20 Devamanyu Hazarika , Roger Zimmermann , Soujanya Poria

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Tuyen Tran , Thao Minh Le , Hung Tran , Truyen Tran

Multi-modal fusion is a fundamental task for the perception of an autonomous driving system, which has recently intrigued many researchers. However, achieving a rather good performance is not an easy task due to the noisy raw data,…

计算机视觉与模式识别 · 计算机科学 2024-12-18 Keli Huang , Botian Shi , Xiang Li , Xin Li , Siyuan Huang , Yikang Li

Transformers have made significant strides across various artificial intelligence domains, including natural language processing, computer vision, and audio processing. This success has naturally garnered considerable interest from both…

图像与视频处理 · 电气工程与系统科学 2024-08-12 Elyas Rashno , Amir Eskandari , Aman Anand , Farhana Zulkernine

Feature matching is a cornerstone task in computer vision, essential for applications such as image retrieval, stereo matching, 3D reconstruction, and SLAM. This survey comprehensively reviews modality-based feature matching, exploring…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Weide Liu , Wei Zhou , Jun Liu , Ping Hu , Jun Cheng , Jungong Han , Weisi Lin

Wearable sensor-based human activity recognition (HAR) has been a research focus in the field of ubiquitous and mobile computing for years. In recent years, many deep models have been applied to HAR problems. However, deep learning methods…

信号处理 · 电气工程与系统科学 2020-12-16 Yujiao Hao , Boyu Wang , Rong Zheng

Multi-modal fusion is a basic task of autonomous driving system perception, which has attracted many scholars' interest in recent years. The current multi-modal fusion methods mainly focus on camera data and LiDAR data, but pay little…

机器人学 · 计算机科学 2022-11-14 Yan Gong , Jianli Lu , Jiayi Wu , Wenzhuo Liu

Human Action Recognition (HAR) aims to understand human behavior and assign a label to each action. It has a wide range of applications, and therefore has been attracting increasing attention in the field of computer vision. Human actions…

计算机视觉与模式识别 · 计算机科学 2022-06-22 Zehua Sun , Qiuhong Ke , Hossein Rahmani , Mohammed Bennamoun , Gang Wang , Jun Liu

The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of intelligent geospatial data interpretation. These models combine various data modalities,…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Ziyue Huang , Hongxi Yan , Qiqi Zhan , Shuai Yang , Mingming Zhang , Chenkai Zhang , YiMing Lei , Zeming Liu , Qingjie Liu , Yunhong Wang