English
Related papers

Related papers: Hierarchical Audio-Visual Information Fusion with …

200 papers

The goal of this work is to enhance balanced multimodal understanding in audio-visual large language models (AV-LLMs) by addressing modality bias without additional training. In current AV-LLMs, audio and video features are typically…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Chaeyoung Jung , Youngjoon Jang , Jongmin Choi , Joon Son Chung

Voice disorders negatively impact the quality of daily life in various ways. However, accurately recognizing the category of pathological features from raw audio remains a considerable challenge due to the limited dataset. A promising…

Sound · Computer Science 2024-10-08 Lipeng Shen , Yifan Xiong , Dongyue Guo , Wei Mo , Lingyu Yu , Hui Yang , Yi Lin

Human multimodal emotion recognition (MER) aims to perceive human emotions via language, visual and acoustic modalities. Despite the impressive performance of previous MER approaches, the inherent multimodal heterogeneities still haunt and…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Yong Li , Yuanzhi Wang , Zhen Cui

Emotion recognition from facial expressions is tremendously useful, especially when coupled with smart devices and wireless multimedia applications. However, the inadequate network bandwidth often limits the spatial resolution of the…

Computer Vision and Pattern Recognition · Computer Science 2017-09-12 Bowen Cheng , Zhangyang Wang , Zhaobin Zhang , Zhu Li , Ding Liu , Jianchao Yang , Shuai Huang , Thomas S. Huang

Audio-visual speech enhancement system is regarded to be one of promising solutions for isolating and enhancing speech of desired speaker. Conventional methods focus on predicting clean speech spectrum via a naive convolution neural network…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-28 Xinmeng Xu , Jianjun Hao

Our experiment adapts several popular deep learning methods as well as some traditional methods on the problem of video emotion recognition. In our experiment, we use the CNN-LSTM architecture for visual information extraction and…

Computer Vision and Pattern Recognition · Computer Science 2017-04-13 Lijie Fan , Yunjie Ke

Although deep learning has yielded impressive performance for face recognition, many studies have shown that different networks learn different feature maps: while some networks are more receptive to pose and illumination others appear to…

Computer Vision and Pattern Recognition · Computer Science 2017-02-16 Navaneeth Bodla , Jingxiao Zheng , Hongyu Xu , Jun-Cheng Chen , Carlos Castillo , Rama Chellappa

We present our preliminary work to determine if patient's vocal acoustic, linguistic, and facial patterns could predict clinical ratings of depression severity, namely Patient Health Questionnaire depression scale (PHQ-8). We proposed a…

Computer Vision and Pattern Recognition · Computer Science 2017-12-01 Aven Samareh , Yan Jin , Zhangyang Wang , Xiangyu Chang , Shuai Huang

Humans are able to comprehend information from multiple domains for e.g. speech, text and visual. With advancement of deep learning technology there has been significant improvement of speech recognition. Recognizing emotion from speech is…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-16 Mandeep Singh , Yuan Fang

Multi-modality image fusion aims at fusing modality-specific (complementarity) and modality-shared (correlation) information from multiple source images. To tackle the problem of the neglect of inter-feature relationships, high-frequency…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Xiaoli Zhang , Liying Wang , Libo Zhao , Xiongfei Li , Siwei Ma

Emotion recognition in conversations is challenging due to the multi-modal nature of the emotion expression. We propose a hierarchical cross-attention model (HCAM) approach to multi-modal emotion recognition using a combination of recurrent…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-10 Soumya Dutta , Sriram Ganapathy

The affective brain-computer interface is a crucial technology for affective interaction and emotional intelligence, emerging as a significant area of research in the human-computer interaction. Compared to single-type features, multi-type…

Human-Computer Interaction · Computer Science 2025-08-11 Xueyuan Xu , Wenjia Dong , Fulin Wei , Li Zhuo

This paper studies deep network architectures to address the problem of video classification. A multi-stream framework is proposed to fully utilize the rich multimodal information in videos. Specifically, we first train three Convolutional…

Computer Vision and Pattern Recognition · Computer Science 2015-11-12 Zuxuan Wu , Yu-Gang Jiang , Xi Wang , Hao Ye , Xiangyang Xue , Jun Wang

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues, demonstrated through cognitive psychology and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shentong Mo , Pedro Morgado

The high feature dimensionality is a challenge in music emotion recognition. There is no common consensus on a relation between audio features and emotion. The MER system uses all available features to recognize emotion; however, this is…

Sound · Computer Science 2022-12-29 Le Cai , Sam Ferguson , Haiyan Lu , Gengfa Fang

The performance of speech emotion recognition (SER) is limited by the insufficient emotion information in unimodal systems and the feature alignment difficulties in multimodal systems. Recently, multimodal large language models (MLLMs) have…

Sound · Computer Science 2025-09-22 Yiqing Yang , Man-Wai Mak

Accurate brain tumor classification is crucial in medical imaging to ensure reliable diagnosis and effective treatment planning. This study introduces a novel double ensembling framework that synergistically combines pre-trained deep…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Zahid Ullah , Jihie Kim

Infrared and visible image fusion, as a hot topic in image processing and image enhancement, aims to produce fused images retaining the detail texture information in visible images and the thermal radiation information in infrared images. A…

Image and Video Processing · Electrical Eng. & Systems 2021-04-15 Zixiang Zhao , Jiangshe Zhang , Shuang Xu , Kai Sun , Chunxia Zhang , Junmin Liu

Multimodal learning has been a popular area of research, yet integrating electroencephalogram (EEG) data poses unique challenges due to its inherent variability and limited availability. In this paper, we introduce a novel multimodal…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Kang Yin , Hye-Bin Shin , Dan Li , Seong-Whan Lee

Multimodal emotion recognition in conversation (MERC) requires representations that effectively integrate signals from multiple modalities. These signals include modality-specific cues, information shared across modalities, and interactions…

Machine Learning · Computer Science 2026-01-22 Anh-Tuan Mai , Cam-Van Thi Nguyen , Duc-Trong Le