English
Related papers

Related papers: Modulation spectral features for speech emotion re…

200 papers

Speech Emotion Recognition (SER) presents a significant yet persistent challenge in human-computer interaction. While deep learning has advanced spoken language processing, achieving high performance on limited datasets remains a critical…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Tai Vu

Speech-based depression detection tools could aid early screening. Here, we propose an interpretable speech foundation model approach to enhance the clinical applicability of such tools. We introduce a speech-level Audio Spectrogram…

Sound · Computer Science 2026-03-26 Qingkun Deng , Saturnino Luz , Sofia de la Fuente Garcia

Deep convolutional neural networks are being actively investigated in a wide range of speech and audio processing applications including speech recognition, audio event detection and computational paralinguistics, owing to their ability to…

Machine Learning · Computer Science 2018-01-16 Che-Wei Huang , Shrikanth. S. Narayanan

Identifying the emotional state from speech is essential for the natural interaction of the machine with the speaker. However, extracting effective features for emotion recognition is difficult, as emotions are ambiguous. We propose a novel…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Dongyang Dai , Zhiyong Wu , Runnan Li , Xixin Wu , Jia Jia , Helen Meng

Intelligent spectrum management is crucial for improving spectrum efficiency and achieving secure utilization of spectrum resources. However, existing intelligent spectrum management methods, typically based on small-scale models, suffer…

Signal Processing · Electrical Eng. & Systems 2025-12-16 Fuhui Zhou , Chunyu Liu , Hao Zhang , Wei Wu , Qihui Wu , Tony Q. S. Quek , Chan-Byoung Chae

This paper addresses the problem of under-determinded speech source separation from multichannel microphone singals, i.e. the convolutive mixtures of multiple sources. The time-domain signals are first transformed to the short-time Fourier…

Sound · Computer Science 2019-04-11 Xiaofei Li , Laurent Girin , Radu Horaud

Speech emotion recognition (SER) remains fragile in real-world conditions because emotional cues are subtle, speaker-dependent, and easily confounded by recording variability, while high-performing deep models typically rely on large and…

Quantum Physics · Physics 2026-05-15 Mahad Mohtashim , Nouhaila Innan , Muhammad Shafique

The performance of single image super-resolution has achieved significant improvement by utilizing deep convolutional neural networks (CNNs). The features in deep CNN contain different types of information which make different contributions…

Computer Vision and Pattern Recognition · Computer Science 2018-10-01 Yanting Hu , Jie Li , Yuanfei Huang , Xinbo Gao

To date, mainstream target speech separation (TSS) approaches are formulated to estimate the complex ratio mask (cRM) of the target speech in time-frequency domain under supervised deep learning framework. However, the existing deep models…

Sound · Computer Science 2021-09-08 Rongzhi Gu , Shi-Xiong Zhang , Yuexian Zou , Dong Yu

Speech emotion recognition (SER) plays a vital role in improving the interactions between humans and machines by inferring human emotion and affective states from speech signals. Whereas recent works primarily focus on mining spatiotemporal…

Sound · Computer Science 2023-10-03 Jiaxin Ye , Xin-cheng Wen , Yujie Wei , Yong Xu , Kunhong Liu , Hongming Shan

Deep Learning models have become potential candidates for auditory neuroscience research, thanks to their recent successes on a variety of auditory tasks. Yet, these models often lack interpretability to fully understand the exact…

Sound · Computer Science 2021-08-04 Rachid Riad , Julien Karadayi , Anne-Catherine Bachoud-Lévi , Emmanuel Dupoux

Automatic emotion recognition plays a key role in computer-human interaction as it has the potential to enrich the next-generation artificial intelligence with emotional intelligence. It finds applications in customer and/or representative…

Sound · Computer Science 2022-02-21 Sarala Padi , Seyed Omid Sadjadi , Dinesh Manocha , Ram D. Sriram

In this paper, we propose a deep learning based system for the task of deepfake audio detection. In particular, the draw input audio is first transformed into various spectrograms using three transformation methods of Short-time Fourier…

Sound · Computer Science 2024-07-03 Lam Pham , Phat Lam , Truong Nguyen , Huyen Nguyen , Alexander Schindler

Recent advancements in transformer-based speech representation models have greatly transformed speech processing. However, there has been limited research conducted on evaluating these models for speech emotion recognition (SER) across…

Computation and Language · Computer Science 2023-08-21 Anant Singh , Akshat Gupta

The high feature dimensionality is a challenge in music emotion recognition. There is no common consensus on a relation between audio features and emotion. The MER system uses all available features to recognize emotion; however, this is…

Sound · Computer Science 2022-12-29 Le Cai , Sam Ferguson , Haiyan Lu , Gengfa Fang

Facial expressions are one of the most powerful ways for depicting specific patterns in human behavior and describing human emotional state. Despite the impressive advances of affective computing over the last decade, automatic video-based…

Computer Vision and Pattern Recognition · Computer Science 2021-01-18 Thomas Teixeira , Eric Granger , Alessandro Lameiras Koerich

Detecting emotions directly from a speech signal plays an important role in effective human-computer interactions. Existing speech emotion recognition models require massive computational and storage resources, making them hard to implement…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-08 Arya Aftab , Alireza Morsali , Shahrokh Ghaemmaghami , Benoit Champagne

Convolutional Neural Networks have been extensively explored in the task of automatic music tagging. The problem can be approached by using either engineered time-frequency features or raw audio as input. Modulation filter bank…

Sound · Computer Science 2021-05-26 Cyrus Vahidi , Charalampos Saitis , György Fazekas

Speech separation has been very successful with deep learning techniques. Substantial effort has been reported based on approaches over spectrogram, which is well known as the standard time-and-frequency cross-domain representation for…

Sound · Computer Science 2019-04-17 Gene-Ping Yang , Chao-I Tuan , Hung-Yi Lee , Lin-shan Lee

Frequency modulation features capture the fine structure of speech formants that constitute beneficial and supplementary to the traditional energy-based cepstral features. Improvements have been demonstrated mainly in GMM-HMM systems for…

Sound · Computer Science 2019-09-04 Isidoros Rodomagoulakis , Petros Maragos