English
Related papers

Related papers: Deep Equilibrium Multimodal Fusion

200 papers

The recent explosion of interest in multimodal applications has resulted in a wide selection of datasets and methods for representing and integrating information from different modalities. Despite these empirical advances, there remain…

Machine learning force fields show great promise in enabling more accurate molecular dynamics simulations compared to manually derived ones. Much of the progress in recent years was driven by exploiting prior knowledge about physical…

Machine Learning · Computer Science 2025-09-11 Andreas Burger , Luca Thiede , Alán Aspuru-Guzik , Nandita Vijaykumar

Multimodal learning aims to imitate human beings to acquire complementary information from multiple modalities for various downstream tasks. However, traditional aggregation-based multimodal fusion methods ignore the inter-modality…

Computer Vision and Pattern Recognition · Computer Science 2023-05-17 Heqing Zou , Meng Shen , Chen Chen , Yuchen Hu , Deepu Rajan , Eng Siong Chng

Multimodal Sentiment Analysis (MSA) stands as a critical research frontier, seeking to comprehensively unravel human emotions by amalgamating text, audio, and visual data. Yet, discerning subtle emotional nuances within audio and video…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Sheng Wu , Xiaobao Wang , Longbiao Wang , Dongxiao He , Jianwu Dang

Multimodal recommender systems leverage diverse data sources, such as user interactions, content features, and contextual information, to address challenges like cold-start and data sparsity. However, existing methods often suffer from one…

Information Retrieval · Computer Science 2026-02-24 Adamya Shyam , Venkateswara Rao Kagita , Bharti Rana , Vikas Kumar

Multimodal information processing has become increasingly important for enhancing image classification performance. However, the intricate and implicit dependencies across different modalities often hinder conventional methods from…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Yang Qiao , Xiaoyu Zhong , Xiaofeng Gu , Zhiguo Yu

Multi-sensor fusion plays a critical role in enhancing perception for autonomous driving, overcoming individual sensor limitations, and enabling comprehensive environmental understanding. This paper first formalizes multi-sensor fusion…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Chuheng Wei , Ziye Qin , Ziyan Zhang , Guoyuan Wu , Matthew J. Barth

The main idea of multimodal recommendation is the rational utilization of the item's multimodal information to improve the recommendation performance. Previous works directly integrate item multimodal features with item ID embeddings,…

Information Retrieval · Computer Science 2023-04-25 Yan Zhou , Jie Guo , Hao Sun , Bin Song , Fei Richard Yu

Visual recognition inside the vehicle cabin leads to safer driving and more intuitive human-vehicle interaction but such systems face substantial obstacles as they need to capture different granularities of driver behaviour while dealing…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Alina Roitberg , Kunyu Peng , Zdravko Marinov , Constantin Seibold , David Schneider , Rainer Stiefelhagen

To overcome the imbalanced multimodal learning problem, where models prefer the training of specific modalities, existing methods propose to control the training of uni-modal encoders from different perspectives, taking the inter-modal…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Yake Wei , Siwei Li , Ruoxuan Feng , Di Hu

Multi-modal image fusion (MMIF) integrates valuable information from different modality images into a fused one. However, the fusion of multiple visible images with different focal regions and infrared images is a unprecedented challenge in…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Xilai Li , Xiaosong Li , Tao Ye , Xiaoqi Cheng , Wuyang Liu , Haishu Tan

The human visual perception system has strong robustness in image fusion. This robustness is based on human visual perception system's characteristics of feature selection and non-linear fusion of different features. In order to simulate…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Aiqing Fang , Xinbo Zhao , Jiaqi Yang , Yanning Zhang

Knowledge Tracing is the process of tracking mastery level of different skills of students for a given learning domain. It is one of the key components for building adaptive learning systems and has been investigated for decades. In…

Machine Learning · Computer Science 2021-11-09 Xinyi Ding , Tao Han , Yili Fang , Eric Larson

Speech emotion recognition (SER) remains a challenging yet crucial task due to the inherent complexity and diversity of human emotions. To address this problem, researchers attempt to fuse information from other modalities via multimodal…

Sound · Computer Science 2024-12-10 Feng Li , Jiusong Luo , Wanjun Xia

Learning holistic computational representations in physical, chemical or biological systems requires the ability to process information from different distributions and modalities within the same model. Thus, the demand for multimodal…

Machine Learning · Computer Science 2025-04-17 Konstantin Hemker , Nikola Simidjievski , Mateja Jamnik

Multimodal semantic segmentation is a pivotal component of computer vision and typically surpasses unimodal methods by utilizing rich information set from various sources.Current models frequently adopt modality-specific frameworks that…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Bingyu Li , Da Zhang , Zhiyuan Zhao , Junyu Gao , Xuelong Li

While VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood. Do these models capture the rich multimodal structures and dynamics from video and text…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Ishaan Singh Rawal , Alexander Matyasko , Shantanu Jaiswal , Basura Fernando , Cheston Tan

Video Question Answering (VideoQA) is a very attractive and challenging research direction aiming to understand complex semantics of heterogeneous data from two domains, i.e., the spatio-temporal video content and the word sequence in…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Chengxiang Yin , Zhengping Che , Kun Wu , Zhiyuan Xu , Qinru Qiu , Jian Tang

Different modalities hold considerable gaps in optimization trajectories, including speeds and paths, which lead to modality laziness and modality clash when jointly training multimodal models, resulting in insufficient and imbalanced…

Machine Learning · Computer Science 2025-06-17 Xiaoyu Ma , Hao Chen , Yongjian Deng

Multimodal medical image fusion plays a crucial role in medical diagnosis by integrating complementary information from different modalities to enhance image readability and clinical applicability. However, existing methods mainly follow…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Haozhe Xiang , Han Zhang , Yu Cheng , Xiongwen Quan , Wanwan Huang