English
Related papers

Related papers: GateFusion: Hierarchical Gated Cross-Modal Fusion …

200 papers

In recent research, slight performance improvement is observed from automatic speech recognition systems to audio-visual speech recognition systems in the end-to-end framework with low-quality videos. Unmatching convergence rates and…

Computation and Language · Computer Science 2024-03-12 Yusheng Dai , Hang Chen , Jun Du , Xiaofei Ding , Ning Ding , Feijun Jiang , Chin-Hui Lee

Early detection of cognitive disorders such as Alzheimer's disease is critical for enabling timely clinical intervention and improving patient outcomes. In this work, we introduce CogniAlign, a multimodal architecture for Alzheimer's…

Machine Learning · Computer Science 2025-10-27 David Ortiz-Perez , Manuel Benavent-Lledo , Javier Rodriguez-Juan , Jose Garcia-Rodriguez , David Tomás

Multimodal learning faces a fundamental tension between deep, fine-grained fusion and computational scalability. While cross-attention models achieve strong performance through exhaustive pairwise fusion, their quadratic complexity is…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Yusuf Shihata

Humans perceive the world by concurrently processing and fusing high-dimensional inputs from multiple modalities such as vision and audio. Machine perception models, in stark contrast, are typically modality-specific and optimised for…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Arsha Nagrani , Shan Yang , Anurag Arnab , Aren Jansen , Cordelia Schmid , Chen Sun

Emotion recognition in multi-speaker conversations faces significant challenges due to speaker ambiguity and severe class imbalance. We propose a novel framework that addresses these issues through three key innovations: (1) a speaker…

Sound · Computer Science 2025-11-19 Xiao Li , Kotaro Funakoshi , Manabu Okumura

Automatic emotion recognition (AER) based on enriched multimodal inputs, including text, speech, and visual clues, is crucial in the development of emotionally intelligent machines. Although complex modality relationships have been proven…

Multimedia · Computer Science 2021-09-16 Shuyun Tang , Zhaojie Luo , Guoshun Nan , Yuichiro Yoshikawa , Ishiguro Hiroshi

Existing robotic manipulation methods primarily rely on visual and proprioceptive observations, which may struggle to infer contact-related interaction states in partially observable real-world environments. Acoustic cues, by contrast,…

Robotics · Computer Science 2026-02-17 Siyuan Li , Jiani Lu , Yu Song , Xianren Li , Bo An , Peng Liu

Most existing sound event detection~(SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts have explored language-driven zero-shot SED by exploiting…

Sound · Computer Science 2025-10-28 Pengfei Cai , Yan Song , Qing Gu , Nan Jiang , Haoyu Song , Ian McLoughlin

Spatial transcriptomics (ST) bridges gene expression and tissue morphology but faces clinical adoption barriers due to technical complexity and prohibitive costs. While computational methods predict gene expression from H&E-stained…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Ziqiao Weng , Yaoyu Fang , Jiahe Qian , Xinkun Wang , Lee AD Cooper , Weidong Cai , Bo Zhou

Overlapped speech detection (OSD) is critical for speech applications in scenario of multi-party conversion. Despite numerous research efforts and progresses, comparing with speech activity detection (VAD), OSD remains an open challenge and…

Sound · Computer Science 2022-09-27 Ziqing Du , Kai Liu , Xucheng Wan , Huan Zhou

Multimodal sentiment analysis is an important research task to predict the sentiment score based on the different modality data from a specific opinion video. Many previous pieces of research have proved the significance of utilizing the…

Computation and Language · Computer Science 2022-08-26 Ming Jiang , Shaoxiong Ji

3D object detection serves as the core basis of the perception tasks in autonomous driving. Recent years have seen the rapid progress of multi-modal fusion strategies for more robust and accurate 3D object detection. However, current…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Bingqi Shen , Shuwei Dai , Yuyin Chen , Rong Xiong , Yue Wang , Yanmei Jiao

Predicting the future behavior of road users is one of the most challenging and important problems in autonomous driving. Applying deep learning to this problem requires fusing heterogeneous world state in the form of rich perception…

Leveraging the synergy of both audio data and visual data is essential for understanding human emotions and behaviors, especially in in-the-wild setting. Traditional methods for integrating such multimodal information often stumble, leading…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Jun Yu , Zerui Zhang , Zhihong Wei , Gongpeng Zhao , Zhongpeng Cai , Yongqi Wang , Guochen Xie , Jichao Zhu , Wangyuan Zhu

Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital…

Sound · Computer Science 2025-10-15 Tianbao Zhang , Jian Zhao , Yuer Li , Zheng Zhu , Ping Hu , Zhaoxin Fan , Wenjun Wu , Xuelong Li

Multi-modal 3D object detectors are dedicated to exploring secure and reliable perception systems for autonomous driving (AD).Although achieving state-of-the-art (SOTA) performance on clean benchmark datasets, they tend to overlook the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Ziying Song , Guoxing Zhang , Lin Liu , Lei Yang , Shaoqing Xu , Caiyan Jia , Feiyang Jia , Li Wang

In this paper, we address three challenges in utterance-level emotion recognition in dialogue systems: (1) the same word can deliver different emotions in different contexts; (2) some emotions are rarely seen in general dialogues; (3)…

Computation and Language · Computer Science 2019-04-10 Wenxiang Jiao , Haiqin Yang , Irwin King , Michael R. Lyu

Multimodal 3D object detection based on deep neural networks has indeed made significant progress. However, it still faces challenges due to the misalignment of scale and spatial information between features extracted from 2D images and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Bonan Ding , Jin Xie , Jing Nie , Jiale Cao

Multimodal emotion recognition often suffers from performance degradation in valence-arousal estimation due to noise and misalignment between audio and visual modalities. To address this challenge, we introduce TAGF, a Time-aware Gated…

Multimedia · Computer Science 2025-07-04 Yubeen Lee , Sangeun Lee , Chaewon Park , Junyeop Cha , Eunil Park

With recent advances in autonomous driving, Voice Control Systems have become increasingly adopted as human-vehicle interaction methods. This technology enables drivers to use voice commands to control the vehicle and will be soon available…

Machine Learning · Computer Science 2021-12-03 Jiwei Guan , Xi Zheng , Chen Wang , Yipeng Zhou , Alireza Jolfa