中文
相关论文

相关论文: Modulation spectral features for speech emotion re…

200 篇论文

Speech Emotion Recognition (SER) presents a significant yet persistent challenge in human-computer interaction. While deep learning has advanced spoken language processing, achieving high performance on limited datasets remains a critical…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Tai Vu

Speech-based depression detection tools could aid early screening. Here, we propose an interpretable speech foundation model approach to enhance the clinical applicability of such tools. We introduce a speech-level Audio Spectrogram…

声音 · 计算机科学 2026-03-26 Qingkun Deng , Saturnino Luz , Sofia de la Fuente Garcia

Deep convolutional neural networks are being actively investigated in a wide range of speech and audio processing applications including speech recognition, audio event detection and computational paralinguistics, owing to their ability to…

机器学习 · 计算机科学 2018-01-16 Che-Wei Huang , Shrikanth. S. Narayanan

Identifying the emotional state from speech is essential for the natural interaction of the machine with the speaker. However, extracting effective features for emotion recognition is difficult, as emotions are ambiguous. We propose a novel…

音频与语音处理 · 电气工程与系统科学 2025-01-03 Dongyang Dai , Zhiyong Wu , Runnan Li , Xixin Wu , Jia Jia , Helen Meng

Intelligent spectrum management is crucial for improving spectrum efficiency and achieving secure utilization of spectrum resources. However, existing intelligent spectrum management methods, typically based on small-scale models, suffer…

信号处理 · 电气工程与系统科学 2025-12-16 Fuhui Zhou , Chunyu Liu , Hao Zhang , Wei Wu , Qihui Wu , Tony Q. S. Quek , Chan-Byoung Chae

This paper addresses the problem of under-determinded speech source separation from multichannel microphone singals, i.e. the convolutive mixtures of multiple sources. The time-domain signals are first transformed to the short-time Fourier…

声音 · 计算机科学 2019-04-11 Xiaofei Li , Laurent Girin , Radu Horaud

Speech emotion recognition (SER) remains fragile in real-world conditions because emotional cues are subtle, speaker-dependent, and easily confounded by recording variability, while high-performing deep models typically rely on large and…

量子物理 · 物理学 2026-05-15 Mahad Mohtashim , Nouhaila Innan , Muhammad Shafique

The performance of single image super-resolution has achieved significant improvement by utilizing deep convolutional neural networks (CNNs). The features in deep CNN contain different types of information which make different contributions…

计算机视觉与模式识别 · 计算机科学 2018-10-01 Yanting Hu , Jie Li , Yuanfei Huang , Xinbo Gao

To date, mainstream target speech separation (TSS) approaches are formulated to estimate the complex ratio mask (cRM) of the target speech in time-frequency domain under supervised deep learning framework. However, the existing deep models…

声音 · 计算机科学 2021-09-08 Rongzhi Gu , Shi-Xiong Zhang , Yuexian Zou , Dong Yu

Speech emotion recognition (SER) plays a vital role in improving the interactions between humans and machines by inferring human emotion and affective states from speech signals. Whereas recent works primarily focus on mining spatiotemporal…

声音 · 计算机科学 2023-10-03 Jiaxin Ye , Xin-cheng Wen , Yujie Wei , Yong Xu , Kunhong Liu , Hongming Shan

Deep Learning models have become potential candidates for auditory neuroscience research, thanks to their recent successes on a variety of auditory tasks. Yet, these models often lack interpretability to fully understand the exact…

声音 · 计算机科学 2021-08-04 Rachid Riad , Julien Karadayi , Anne-Catherine Bachoud-Lévi , Emmanuel Dupoux

Automatic emotion recognition plays a key role in computer-human interaction as it has the potential to enrich the next-generation artificial intelligence with emotional intelligence. It finds applications in customer and/or representative…

声音 · 计算机科学 2022-02-21 Sarala Padi , Seyed Omid Sadjadi , Dinesh Manocha , Ram D. Sriram

In this paper, we propose a deep learning based system for the task of deepfake audio detection. In particular, the draw input audio is first transformed into various spectrograms using three transformation methods of Short-time Fourier…

声音 · 计算机科学 2024-07-03 Lam Pham , Phat Lam , Truong Nguyen , Huyen Nguyen , Alexander Schindler

Recent advancements in transformer-based speech representation models have greatly transformed speech processing. However, there has been limited research conducted on evaluating these models for speech emotion recognition (SER) across…

计算与语言 · 计算机科学 2023-08-21 Anant Singh , Akshat Gupta

The high feature dimensionality is a challenge in music emotion recognition. There is no common consensus on a relation between audio features and emotion. The MER system uses all available features to recognize emotion; however, this is…

声音 · 计算机科学 2022-12-29 Le Cai , Sam Ferguson , Haiyan Lu , Gengfa Fang

Facial expressions are one of the most powerful ways for depicting specific patterns in human behavior and describing human emotional state. Despite the impressive advances of affective computing over the last decade, automatic video-based…

计算机视觉与模式识别 · 计算机科学 2021-01-18 Thomas Teixeira , Eric Granger , Alessandro Lameiras Koerich

Detecting emotions directly from a speech signal plays an important role in effective human-computer interactions. Existing speech emotion recognition models require massive computational and storage resources, making them hard to implement…

音频与语音处理 · 电气工程与系统科学 2021-10-08 Arya Aftab , Alireza Morsali , Shahrokh Ghaemmaghami , Benoit Champagne

Convolutional Neural Networks have been extensively explored in the task of automatic music tagging. The problem can be approached by using either engineered time-frequency features or raw audio as input. Modulation filter bank…

声音 · 计算机科学 2021-05-26 Cyrus Vahidi , Charalampos Saitis , György Fazekas

Speech separation has been very successful with deep learning techniques. Substantial effort has been reported based on approaches over spectrogram, which is well known as the standard time-and-frequency cross-domain representation for…

声音 · 计算机科学 2019-04-17 Gene-Ping Yang , Chao-I Tuan , Hung-Yi Lee , Lin-shan Lee

Frequency modulation features capture the fine structure of speech formants that constitute beneficial and supplementary to the traditional energy-based cepstral features. Improvements have been demonstrated mainly in GMM-HMM systems for…

声音 · 计算机科学 2019-09-04 Isidoros Rodomagoulakis , Petros Maragos