English
Related papers

Related papers: Evaluating Gammatone Frequency Cepstral Coefficien…

200 papers

Speech emotion recognition plays a crucial role in human-machine interaction systems. Recently various optimized Transformers have been successfully applied to speech emotion recognition. However, the existing Transformer architectures…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-02 Zhipeng Li , Xiaofen Xing , Yuanbo Fang , Weibin Zhang , Hengsheng Fan , Xiangmin Xu

Ultra-low-bitrate speech coding is pivotal for bandwidth-constrained communication and deep compression, yet maintaining naturalness and speaker identity at such extreme bit budgets remains challenging due to pronounced information loss and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-26 Hui-Peng Du , Yang Ai , Xiao-Hang Jiang , Yuan Tian , Zhen-Hua Ling

In this paper, we propose a new methodology for emotional speech recognition using visual deep neural network models. We employ the transfer learning capabilities of the pre-trained computer vision deep models to have a mandate for the…

Computer Vision and Pattern Recognition · Computer Science 2022-04-08 Waleed Ragheb , Mehdi Mirzapour , Ali Delfardi , Hélène Jacquenet , Lawrence Carbon

In this work, a sentiment analysis method that is capable of accepting audio of any length, without being fixed a priori, is proposed. Mel spectrogram and Mel Frequency Cepstral Coefficients are used as audio description methods and a Fully…

Speech deepfake detection (DFD) has benefited from diverse acoustic and semantic speech representations, many of which encode valuable speech information and are costly to train. Existing approaches typically enhance DFD by tuning the…

Sound · Computer Science 2026-02-26 Yupei Li , Chenyang Lyu , Longyue Wang , Weihua Luo , Kaifu Zhang , Björn W. Schuller

Speech Emotion Recognition (SER) plays a pivotal role in enhancing human-computer interaction by enabling a deeper understanding of emotional states across a wide range of applications, contributing to more empathetic and effective…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-25 Amirali Soltani Tehrani , Niloufar Faridani , Ramin Toosi

Speech emotion recognition systems have high prediction latency because of the high computational requirements for deep learning models and low generalizability mainly because of the poor reliability of emotional measurements across…

Sound · Computer Science 2023-02-23 Abdul Rehman , Zhen-Tao Liu , Min Wu , Wei-Hua Cao , Cheng-Shan Jiang

Voice Activity Detection (VAD) plays a key role in speech processing, often utilizing hand-crafted or neural features. This study examines the effectiveness of Mel-Frequency Cepstral Coefficients (MFCCs) and pre-trained model (PTM)…

Sound · Computer Science 2025-06-03 Kumud Tripathi , Chowdam Venkata Kumar , Pankaj Wasnik

Recently, attention mechanisms have been applied successfully in neural network-based speaker verification systems. Incorporating the Squeeze-and-Excitation block into convolutional neural networks has achieved remarkable performance.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-12 Mufan Sang , John H. L. Hansen

With the development of computer -systems that can collect and analyze enormous volumes of data, the medical profession is establishing several non-invasive tools. This work attempts to develop a non-invasive technique for identifying…

Sound · Computer Science 2023-03-16 Hafsa Gulzar , Jiyun Li , Arslan Manzoor , Sadaf Rehmat , Usman Amjad , Hadiqa Jalil Khan

Recent successful applications of convolutional neural networks (CNNs) to audio classification and speech recognition have motivated the search for better input representations for more efficient training. Visual displays of an audio…

Computer Vision and Pattern Recognition · Computer Science 2017-06-23 M. Huzaifah

Fast Fourier convolution (FFC) is the recently proposed neural operator showing promising performance in several computer vision problems. The FFC operator allows employing large receptive field operations within early layers of the neural…

Sound · Computer Science 2022-04-08 Ivan Shchekotov , Pavel Andreev , Oleg Ivanov , Aibek Alanov , Dmitry Vetrov

Spectrogram is commonly used as the input feature of deep neural networks to learn the high(er)-level time-frequency pattern of speech signal for speech emotion recognition (SER). \textcolor{black}{Generally, different emotions correspond…

Sound · Computer Science 2022-10-25 Cheng Lu , Wenming Zheng , Hailun Lian , Yuan Zong , Chuangao Tang , Sunan Li , Yan Zhao

Emotion detection is a central problem in NLP, with recent progress driven by transformer-based models trained on established datasets. However, little is known about the linguistic regularities that characterize how emotions are expressed…

Computation and Language · Computer Science 2026-03-24 Florian Lecourt , Madalina Croitoru , Konstantin Todorov

Acoustic scene classification is a process of characterizing and classifying the environments from sound recordings. The first step is to generate features (representations) from the recorded sound and then classify the background…

The emotion detection technology to enhance human decision-making is an important research issue for real-world applications, but real-life emotion datasets are relatively rare and small. The experiments conducted in this paper use the…

Computation and Language · Computer Science 2023-06-13 Théo Deschamps-Berger , Lori Lamel , Laurence Devillers

Emotion recognition is essential across numerous fields, including medical applications and brain-computer interface (BCI). Emotional responses include behavioral reactions, such as tone of voice and body movement, and changes in…

Signal Processing · Electrical Eng. & Systems 2024-10-02 Eleonora Lopez , Aurelio Uncini , Danilo Comminiello

In this paper, we propose a novel deep inductive transfer learning framework, named feature distribution adaptation network, to tackle the challenging multi-modal speech emotion recognition problem. Our method aims to use deep transfer…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Shaokai Li , Yixuan Ji , Peng Song , Haoqin Sun , Wenming Zheng

Speech Emotion Recognition (SER) aims to help the machine to understand human's subjective emotion from only audio information. However, extracting and utilizing comprehensive in-depth audio information is still a challenging task. In this…

Sound · Computer Science 2022-03-30 Heqing Zou , Yuke Si , Chen Chen , Deepu Rajan , Eng Siong Chng

Environmental sound classification (ESC) has gained significant attention due to its diverse applications in smart city monitoring, fault detection, acoustic surveillance, and manufacturing quality control. To enhance CNN performance,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-25 Parinaz Binandeh Dehaghania , Danilo Penab , A. Pedro Aguiar