English
Related papers

Related papers: Graph-based multi-Feature fusion method for speech…

200 papers

Semantic vector embedding techniques have proven useful in learning semantic representations of data across multiple domains. A key application enabled by such techniques is the ability to measure semantic similarity between given data…

Computation and Language · Computer Science 2020-09-01 Shalisha Witherspoon , Dean Steuer , Graham Bent , Nirmit Desai

Embeddings in AI convert symbolic structures into fixed-dimensional vectors, effectively fusing multiple signals. However, the nature of this fusion in real-world data is often unclear. To address this, we introduce two methods: (1)…

Machine Learning · Computer Science 2023-11-21 Zhijin Guo , Zhaozhen Xu , Martha Lewis , Nello Cristianini

Speech deepfake detection (DFD) has benefited from diverse acoustic and semantic speech representations, many of which encode valuable speech information and are costly to train. Existing approaches typically enhance DFD by tuning the…

Sound · Computer Science 2026-02-26 Yupei Li , Chenyang Lyu , Longyue Wang , Weihua Luo , Kaifu Zhang , Björn W. Schuller

In recent years, deep learning models have demonstrated remarkable success in various domains, such as computer vision, natural language processing, and speech recognition. However, the generalization capabilities of these models can be…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Neelesh Mungoli

We introduce PGF-Net (Progressive Gated-Fusion Network), a novel deep learning framework designed for efficient and interpretable multimodal sentiment analysis. Our framework incorporates three primary innovations. Firstly, we propose a…

Machine Learning · Computer Science 2025-08-25 Bin Wen , Tien-Ping Tan

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

This paper presents the results of the SUN team for the Compound Expressions Recognition Challenge of the 6th ABAW Competition. We propose a novel audio-visual method for compound expression recognition. Our method relies on emotion…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Elena Ryumina , Maxim Markitantov , Dmitry Ryumin , Heysem Kaya , Alexey Karpov

This paper presents a novel deep neural network (DNN) for multimodal fusion of audio, video and text modalities for emotion recognition. The proposed DNN architecture has independent and shared layers which aim to learn the representation…

Computer Vision and Pattern Recognition · Computer Science 2019-07-09 Juan D. S. Ortega , Mohammed Senoussaoui , Eric Granger , Marco Pedersoli , Patrick Cardinal , Alessandro L. Koerich

Emotional Recognition in Conversation (ERC) is valuable for diagnosing health conditions such as autism and depression, and for understanding the emotions of individuals who struggle to express their feelings. Current ERC methods primarily…

Human-Computer Interaction · Computer Science 2026-05-06 Zijian Kang , Yueyang Li , Shengyu Gong , Weiming Zeng , Hongjie Yan , Lingbin Bian , Zhiguo Zhang , Wai Ting Siok , Nizhuan Wang

Multimodal learning involves integrating information from various modalities to enhance learning and comprehension. We compare three modality fusion strategies in person identification and verification by processing two modalities: voice…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-05 Aref Farhadipour , Masoumeh Chapariniya , Teodora Vukovic , Volker Dellwo

Recognizing a face based on its attributes is an easy task for a human to perform as it is a cognitive process. In recent years, Face Recognition is achieved with different kinds of facial features which were used separately or in a…

Computer Vision and Pattern Recognition · Computer Science 2010-11-10 S. Sakthivel , R. Lakshmipathi

Cognitive Behavioral Therapy (CBT) is a goal-oriented psychotherapy for mental health concerns implemented in a conversational setting with broad empirical support for its effectiveness across a range of presenting problems and client…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-16 Zhuohao Chen , Nikolaos Flemotomos , Victor Ardulov , Torrey A. Creed , Zac E. Imel , David C. Atkins , Shrikanth Narayanan

Multimodal emotion understanding requires effective integration of text, audio, and visual modalities for both discrete emotion recognition and continuous sentiment analysis. We present EGMF, a unified framework combining expert-guided…

Computation and Language · Computer Science 2026-01-13 Jiaqi Qiao , Xiujuan Xu , Xinran Li , Yu Liu

In this paper we explore audiovisual emotion recognition under noisy acoustic conditions with a focus on speech features. We attempt to answer the following research questions: (i) How does speech emotion recognition perform on noisy data?…

Sound · Computer Science 2021-03-03 Michael Neumann , Ngoc Thang Vu

We propose a multi-channel speech enhancement approach with a novel two-stage feature fusion method and a pre-trained acoustic model in a multi-task learning paradigm. In the first fusion stage, the time-domain and frequency-domain features…

Sound · Computer Science 2021-09-27 Quandong Wang , Junnan Wu , Zhao Yan , Sichong Qian , Liyong Guo , Lichun Fan , Weiji Zhuang , Peng Gao , Yujun Wang

Emotion plays a fundamental role in human interaction, and therefore systems capable of identifying emotions in speech are crucial in the context of human-computer interaction. Speech emotion recognition (SER) is a challenging problem,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Lucas Ueda , João Lima , Leonardo Marques , Paula Costa

EEG based multi-dimension emotion recognition has attracted substantial research interest in human computer interfaces. However, the high dimensionality of EEG features, coupled with limited sample sizes, frequently leads to classifier…

Human-Computer Interaction · Computer Science 2025-08-08 Tianze Yu , Junming Zhang , Wenjia Dong , Xueyuan Xu , Li Zhuo

Speech emotion recognition (SER) is crucial for enhancing affective computing and enriching the domain of human-computer interaction. However, the main challenge in SER lies in selecting relevant feature representations from speech signals…

Sound · Computer Science 2024-12-16 Niloy Kumar Kundu , Sarah Kobir , Md. Rayhan Ahmed , Tahmina Aktar , Niloya Roy

In this paper, we propose MMER, a novel Multimodal Multi-task learning approach for Speech Emotion Recognition. MMER leverages a novel multimodal network based on early-fusion and cross-modal self-attention between text and acoustic…

Computation and Language · Computer Science 2023-06-06 Sreyan Ghosh , Utkarsh Tyagi , S Ramaneswaran , Harshvardhan Srivastava , Dinesh Manocha

Emotion recognition in multi-speaker conversations faces significant challenges due to speaker ambiguity and severe class imbalance. We propose a novel framework that addresses these issues through three key innovations: (1) a speaker…

Sound · Computer Science 2025-11-19 Xiao Li , Kotaro Funakoshi , Manabu Okumura