English
Related papers

Related papers: TMDC: A Two-Stage Modality Denoising and Complemen…

200 papers

In this paper, we consider the problem of multimodal data analysis with a use case of audiovisual emotion recognition. We propose an architecture capable of learning from raw data and describe three variants of it with distinct modality…

Computer Vision and Pattern Recognition · Computer Science 2022-01-27 Kateryna Chumachenko , Alexandros Iosifidis , Moncef Gabbouj

Multimodal emotion recognition leverages complementary information across modalities to gain performance. However, we cannot guarantee that the data of all modalities are always present in practice. In the studies to predict the missing…

Computer Vision and Pattern Recognition · Computer Science 2022-10-28 Haolin Zuo , Rui Liu , Jinming Zhao , Guanglai Gao , Haizhou Li

Incomplete scenario is a prevalent, practical, yet challenging setting in Multimodal Recommendations (MMRec), where some item modalities are missing due to various factors. Recently, a few efforts have sought to improve the recommendation…

Information Retrieval · Computer Science 2025-02-25 Jin Li , Shoujin Wang , Qi Zhang , Shui Yu , Fang Chen

Multimodal large language models (MLLMs) promise enhanced reasoning by integrating diverse inputs such as text, vision, and audio. Yet cross-modal reasoning remains underexplored, with conflicting reports on whether added modalities help or…

Computation and Language · Computer Science 2026-05-01 Yucheng Wang , Yifan Hou , Aydin Javadov , Mubashara Akhtar , Mrinmaya Sachan

Multimodal learning has exhibited a significant advantage in affective analysis tasks owing to the comprehensive information of various modalities, particularly the complementary information. Thus, many emerging studies focus on…

Artificial Intelligence · Computer Science 2024-04-09 Ying Zhou , Xuefeng Liang , Han Chen , Yin Zhao , Xin Chen , Lida Yu

Detecting emotional inconsistency across modalities is a key challenge in affective computing, as speech and text often convey conflicting cues. Existing approaches generally rely on incomplete emotion representations and employ…

Multimedia · Computer Science 2025-09-25 Zongyi Li , Junchuan Zhao , Francis Bu Sung Lee , Andrew Zi Han Yee

Multimodal learning with incomplete input data (missing modality) is practical and challenging. In this work, we conduct an in-depth analysis of this challenge and find that modality dominance has a significant negative impact on the model…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Hao Wang , Shengda Luo , Guosheng Hu , Jianguo Zhang

Multimodal learning typically relies on the assumption that all modalities are fully available during both the training and inference phases. However, in real-world scenarios, consistently acquiring complete multimodal data presents…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Donggeun Kim , Taesup Kim

Recent progress has been made in detecting early stage dementia entirely through recordings of patient speech. Multimodal speech analysis methods were applied to the PROCESS challenge, which requires participants to use audio recordings of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-14 Lei Chi , Arav Sharma , Ari Gebhardt , Joseph T. Colonel

Accurate extraction of molecular representations is a critical step in the drug discovery process. In recent years, significant progress has been made in molecular representation learning methods, among which multi-modal molecular…

Machine Learning · Computer Science 2025-05-13 Rong Yin , Ruyue Liu , Xiaoshuai Hao , Xingrui Zhou , Yong Liu , Can Ma , Weiping Wang

Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in performance and may even be critical in certain cases, they come…

Sound · Computer Science 2025-01-31 Joanna Hong , Sanjeel Parekh , Honglie Chen , Jacob Donley , Ke Tan , Buye Xu , Anurag Kumar

Missing modalities cause severe failures in multimodal recommender systems. User histories, item text, and visual evidence are frequently absent during cold-start scenarios, exactly when recommendation quality matters most. Existing…

Information Retrieval · Computer Science 2026-05-26 Jinze Wang , Yangchen Zeng , Tiehua Zhang , Lu Zhang , Yuze Liu , Zhishu Shen , Jiong Jin , Zhu Sun

We aim to develop a fundamental understanding of modality collapse, a recently observed empirical phenomenon wherein models trained for multimodal fusion tend to rely only on a subset of the modalities, ignoring the rest. We show that…

Machine Learning · Computer Science 2025-08-18 Abhra Chaudhuri , Anjan Dutta , Tu Bui , Serban Georgescu

Multimodal sentiment analysis is an important research task to predict the sentiment score based on the different modality data from a specific opinion video. Many previous pieces of research have proved the significance of utilizing the…

Computation and Language · Computer Science 2022-08-26 Ming Jiang , Shaoxiong Ji

Missing modalities remain a major challenge for multimodal sensing, because most existing methods adapt the fusion process to the observed subset by dropping absent branches, using subset-specific fusion, or reconstructing missing features.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Hao Wang , Yanyu Qian , Pengcheng Weng , Zixuan Xia , William Dan , Yangxin Xu , Fei Wang

Optimizing multimodal waveguide performance depends on modal analysis; however, existing methods focus predominantly on modal power distribution (MPD) and, limited by experimental hardware and conditions, exhibit low accuracy, poor…

Computational Physics · Physics 2025-11-06 Jingtong Li , Dongting Huang , Minhui Xiong , Mingzhi Li

With the increasing popularity of video sharing websites such as YouTube and Facebook, multimodal sentiment analysis has received increasing attention from the scientific community. Contrary to previous works in multimodal sentiment…

Machine Learning · Computer Science 2018-02-06 Minghai Chen , Sen Wang , Paul Pu Liang , Tadas Baltrušaitis , Amir Zadeh , Louis-Philippe Morency

Multimodal learning, which integrates diverse data sources such as images, text, and structured data, has proven superior to unimodal counterparts in high-stakes decision-making. However, while performance gains remain the gold standard for…

Artificial Intelligence · Computer Science 2025-05-07 Kishore Sampath , Pratheesh , Ayaazuddin Mohammad , Resmi Ramachandranpillai

The study of human emotions, traditionally a cornerstone in fields like psychology and neuroscience, has been profoundly impacted by the advent of artificial intelligence (AI). Multiple channels, such as speech (voice) and facial…

Speech data collected in real-world scenarios often encounters two issues. First, multiple sources may exist simultaneously, and the number of sources may vary with time. Second, the existence of background noise in recording is inevitable.…

Sound · Computer Science 2020-05-21 Yuan-Kuei Wu , Chao-I Tuan , Hung-yi Lee , Yu Tsao