中文
相关论文

相关论文: Enhancing Speech Emotion Recognition using Dynamic…

200 篇论文

Speech emotion recognition is a crucial problem manifesting in a multitude of applications such as human computer interaction and education. Although several advancements have been made in the recent years, especially with the advent of…

声音 · 计算机科学 2021-03-05 Panagiotis Tzirakis , Anh Nguyen , Stefanos Zafeiriou , Björn W. Schuller

Emotion recognition plays a crucial role in various domains of human-robot interaction. In long-term interactions with humans, robots need to respond continuously and accurately, however, the mainstream emotion recognition methods mostly…

人机交互 · 计算机科学 2024-01-23 Zihan Lin , Francisco Cruz , Eduardo Benitez Sandoval

The process of identifying human emotion and affective states from speech is known as speech emotion recognition (SER). This is based on the observation that tone and pitch in the voice frequently convey underlying emotion. Speech…

声音 · 计算机科学 2024-06-18 Nishargo Nigar

Foundation models have shown superior performance for speech emotion recognition (SER). However, given the limited data in emotion corpora, finetuning all parameters of large pre-trained models for SER can be both resource-intensive and…

音频与语音处理 · 电气工程与系统科学 2024-04-02 Nineli Lashkarashvili , Wen Wu , Guangzhi Sun , Philip C. Woodland

The use of data assimilation for the merging of observed data with dynamical models is becoming standard in modern physics. If a parametric model is known, methods such as Kalman filtering have been developed for this purpose. If no model…

数据分析、统计与概率 · 物理学 2018-01-17 Franz Hamilton , Tyrus Berry , Timothy Sauer

Speech Emotion Recognition is a crucial area of research in human-computer interaction. While significant work has been done in this field, many state-of-the-art networks struggle to accurately recognize emotions in speech when the data is…

音频与语音处理 · 电气工程与系统科学 2025-01-23 Rashedul Hasan , Meher Nigar , Nursadul Mamun , Sayan Paul

Next to decision tree and k-nearest neighbours algorithms deep convolutional neural networks (CNNs) are widely used to classify audio data in many domains like music, speech or environmental sounds. To train a specific CNN various spectral…

声音 · 计算机科学 2025-09-16 Friedrich Wolf-Monheim

This paper explores the problem of training a recurrent neural network from noisy data. While neural network based dynamic predictors perform well with noise-free training data, prediction with noisy inputs during training phase poses a…

系统与控制 · 电气工程与系统科学 2023-04-04 Debdipta Goswami

In Speech Emotion Recognition (SER), textual data is often used alongside audio signals to address their inherent variability. However, the reliance on human annotated text in most research hinders the development of practical SER systems.…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Yuanchao Li , Zeyu Zhao , Ondrej Klejch , Peter Bell , Catherine Lai

Despite the widespread utilization of deep neural networks (DNNs) for speech emotion recognition (SER), they are severely restricted due to the paucity of labeled data for training. Recently, segment-based approaches for SER have been…

音频与语音处理 · 电气工程与系统科学 2021-03-31 Shuiyang Mao , P. C. Ching , Tan Lee

Recently, it has been demonstrated experimentally that adaptive estimation of a continuously varying optical phase provides superior accuracy in the phase estimate compared to static estimation. Here, we show that the mean-square error in…

量子物理 · 物理学 2012-12-12 Shibdas Roy , Ian R. Petersen , Elanor H. Huntington

In this paper, we study different approaches for classifying emotions from speech using acoustic and text-based features. We propose to obtain contextualized word embeddings with BERT to represent the information contained in speech…

机器学习 · 计算机科学 2024-03-28 Leonardo Pepino , Pablo Riera , Luciana Ferrer , Agustin Gravano

Label smoothing is a widely used technique in various domains, such as text classification, image classification and speech recognition, known for effectively combating model overfitting. However, there is little fine-grained analysis on…

计算与语言 · 计算机科学 2024-02-26 Yijie Gao , Shijing Si , Hua Luo , Haixia Sun , Yugui Zhang

In this study, we address the challenges associated with accurately determining gaze location on a screen, which is often compromised by noise from factors such as eye tracker limitations, calibration drift, ambient lighting changes, and…

数值分析 · 数学 2025-04-21 Thoa Thieu , Roderick Melnik

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application…

声音 · 计算机科学 2024-12-18 Jinyi Mi , Xiaohan Shi , Ding Ma , Jiajun He , Takuya Fujimura , Tomoki Toda

Recent research has focused on applying speech large language model (SLLM) to improve speech emotion recognition (SER). However, the inherently high frame rate in speech modality severely limits the signal processing and understanding…

计算与语言 · 计算机科学 2025-09-25 Jialong Mai , Xiaofen Xing , Yawei Li , Weidong Chen , Zhipeng Li , Jingyuan Xing , Xiangmin Xu

Speech emotion recognition (SER), the task of identifying the expression of emotion from spoken content, is challenging due to the difficulty in extracting representations that capture emotional attributes from speech. The scarcity of…

音频与语音处理 · 电气工程与系统科学 2025-08-27 Soumya Dutta , Sriram Ganapathy

This work explores the use of constant-Q transform based modulation spectral features (CQT-MSF) for speech emotion recognition (SER). The human perception and analysis of sound comprise of two important cognitive parts: early auditory…

音频与语音处理 · 电气工程与系统科学 2023-01-18 Premjeet Singh , Md Sahidullah , Goutam Saha

Emotion embedding space learned from references is a straightforward approach for emotion transfer in encoder-decoder structured emotional text to speech (TTS) systems. However, the transferred emotion in the synthetic speech is not…

声音 · 计算机科学 2020-11-18 Tao Li , Shan Yang , Liumeng Xue , Lei Xie

Speech emotion recognition (SER) has advanced significantly for the sake of deep-learning methods, while textual information further enhances its performance. However, few studies have focused on the physiological information during speech…

声音 · 计算机科学 2025-11-12 Ziqian Zhang , Min Huang , Zhongzhe Xiao