English
Related papers

Related papers: Everything at Once -- Multi-modal Fusion Transform…

200 papers

Human emotion can be presented in different modes i.e., audio, video, and text. However, the contribution of each mode in exhibiting each emotion is not uniform. Furthermore, the availability of complete mode-specific details may not always…

Artificial Intelligence · Computer Science 2024-02-20 Naresh Kumar Devulapally , Sidharth Anand , Sreyasee Das Bhattacharjee , Junsong Yuan

Combining complementary information from multiple modalities is intuitively appealing for improving the performance of learning-based approaches. However, it is challenging to fully leverage different modalities due to practical challenges…

Machine Learning · Statistics 2018-05-31 Kuan Liu , Yanen Li , Ning Xu , Prem Natarajan

Utilizing the sensor characteristics of the audio, visible camera, and thermal camera, the robustness of person recognition can be enhanced. Existing multimodal person recognition frameworks are primarily formulated assuming that multimodal…

Multimedia · Computer Science 2022-10-25 Vijay John , Yasutomo Kawanishi

Multimodal retrieval is the task of aggregating information from queries across heterogeneous modalities to retrieve desired targets. State-of-the-art multimodal retrieval models can understand complex queries, yet they are typically…

Information Retrieval · Computer Science 2026-03-25 Chuong Huynh , Manh Luong , Abhinav Shrivastava

Cross-modal alignment Learning integrates information from different modalities like text, image, audio and video to create unified models. This approach develops shared representations and learns correlations between modalities, enabling…

Computer Vision and Pattern Recognition · Computer Science 2024-09-19 Bilal Faye , Hanane Azzag , Mustapha Lebbah

Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in performance and may even be critical in certain cases, they come…

Sound · Computer Science 2025-01-31 Joanna Hong , Sanjeel Parekh , Honglie Chen , Jacob Donley , Ke Tan , Buye Xu , Anurag Kumar

Multimodal learning has been lacking principled ways of combining information from different modalities and learning a low-dimensional manifold of meaningful representations. We study multimodal learning and sensor fusion from a latent…

Machine Learning · Computer Science 2019-04-24 Lijiang Guo

In this study, we propose a method for jointly learning of images and videos using a single model. In general, images and videos are often trained by separate models. We propose in this paper a method that takes a batch of images as input…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Shuki Shimizu , Toru Tamaki

Understanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Dingkang Yang , Mingcheng Li , Linhao Qu , Kun Yang , Peng Zhai , Song Wang , Lihua Zhang

Multimodal federated learning in real-world settings often encounters incomplete and heterogeneous data across clients. This results in misaligned local feature representations that limit the effectiveness of model aggregation. Unlike prior…

Machine Learning · Computer Science 2025-10-28 Duong M. Nguyen , Trong Nghia Hoang , Thanh Trung Huynh , Quoc Viet Hung Nguyen , Phi Le Nguyen

Most existing speech disfluency detection techniques only rely upon acoustic data. In this work, we present a practical multimodal disfluency detection approach that leverages available video data together with audio. We curate an…

Computation and Language · Computer Science 2024-06-12 Payal Mohapatra , Shamika Likhite , Subrata Biswas , Bashima Islam , Qi Zhu

Target speech separation refers to extracting a target speaker's voice from an overlapped audio of simultaneous talkers. Previously the use of visual modality for target speech separation has demonstrated great potentials. This work…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-26 Rongzhi Gu , Shi-Xiong Zhang , Yong Xu , Lianwu Chen , Yuexian Zou , Dong Yu

Representation learning for sketch-based image retrieval has mostly been tackled by learning embeddings that discard modality-specific information. As instances from different modalities can often provide complementary information…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Abhra Chaudhuri , Massimiliano Mancini , Yanbei Chen , Zeynep Akata , Anjan Dutta

In the real world, where information is abundant and diverse across different modalities, understanding and utilizing various data types to improve retrieval systems is a key focus of research. Multimodal composite retrieval integrates…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Suyan Li , Fuxiang Huang , Lei Zhang

Traditional multimodal learners find unified representations for tasks like visual question answering, but rely heavily on paired datasets. However, an overlooked yet potentially powerful question is: can one leverage auxiliary unpaired…

Machine Learning · Computer Science 2025-10-10 Sharut Gupta , Shobhita Sundaram , Chenyu Wang , Stefanie Jegelka , Phillip Isola

We present a novel multimodal multitask network and associated training algorithm. The method is capable of ingesting data from approximately 12 different modalities namely image, video, audio, text, depth, point cloud, time series,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Siddharth Srivastava , Gaurav Sharma

The construction of models for video action classification progresses rapidly. However, the performance of those models can still be easily improved by ensembling with the same models trained on different modalities (e.g. Optical flow).…

Computer Vision and Pattern Recognition · Computer Science 2020-11-06 Stepan Komkov , Maksim Dzabraev , Aleksandr Petiushko

Multimodal data fusion is essential for applications requiring the integration of diverse data sources, especially in the presence of incomplete or sparsely available modalities. This paper presents a comparative study of three multimodal…

Machine Learning · Computer Science 2025-01-03 Josiah Bjorgaard

The ability to quickly learn a new task with minimal instruction - known as few-shot learning - is a central aspect of intelligent agents. Classical few-shot benchmarks make use of few-shot samples from a single modality, but such samples…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Zhiqiu Lin , Samuel Yu , Zhiyi Kuang , Deepak Pathak , Deva Ramanan

Multimodal semantic segmentation shows significant potential for enhancing segmentation accuracy in complex scenes. However, current methods often incorporate specialized feature fusion modules tailored to specific modalities, thereby…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Bingyu Li , Da Zhang , Zhiyuan Zhao , Junyu Gao , Xuelong Li