中文
相关论文

相关论文: Enhancing Speech Emotion Recognition using Dynamic…

200 篇论文

This paper proposes a deep denoising auto-encoder technique to extract better acoustic features for speech synthesis. The technique allows us to automatically extract low-dimensional features from high dimensional spectral features in a…

声音 · 计算机科学 2015-06-18 Zhenzhou Wu , Shinji Takaki , Junichi Yamagishi

Acoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion…

计算与语言 · 计算机科学 2018-04-02 Egor Lakomkin , Cornelius Weber , Sven Magg , Stefan Wermter

To improve the performance of speaker identification systems, an effective and robust method is proposed to extract speech features, capable of operating in noisy environment. Based on the time-frequency multi-resolution property of wavelet…

声音 · 计算机科学 2010-03-31 Mahmoud I. Abdalla , Hanaa S. Ali

Multimodal emotion recognition identifies human emotions from various data modalities like video, text, and audio. However, we found that this task can be easily affected by noisy information that does not contain useful semantics. To this…

多媒体 · 计算机科学 2023-05-05 Yuanyuan Liu , Haoyu Zhang , Yibing Zhan , Zijing Chen , Guanghao Yin , Lin Wei , Zhe Chen

In this paper we explore audiovisual emotion recognition under noisy acoustic conditions with a focus on speech features. We attempt to answer the following research questions: (i) How does speech emotion recognition perform on noisy data?…

声音 · 计算机科学 2021-03-03 Michael Neumann , Ngoc Thang Vu

Feature mapping using deep neural networks is an effective approach for single-channel speech enhancement. Noisy features are transformed to the enhanced ones through a mapping network and the mean square errors between the enhanced and…

音频与语音处理 · 电气工程与系统科学 2019-05-02 Zhong Meng , Jinyu Li , Yifan Gong , Biing-Hwang , Juang

Nanomechanical resonant sensors are used in mass spectrometry via detection of resonance frequency jumps. There is a fundamental trade-off between detection speed and accuracy. Temporal and size resolution are limited by the resonator…

仪器与探测器 · 物理学 2024-01-17 Mete Erdogan , Nuri Berke Baytekin , Serhat Emre Coban , Alper Demir

In this work, we consider a sensor selection drawn at random by a sampling with replacement policy for a linear time-invariant dynamical system subject to process and measurement noise. We employ the Kalman filter to estimate the state of…

系统与控制 · 电气工程与系统科学 2023-03-15 Christopher I. Calle , Shaunak D. Bopardikar

Advancements in spoken language processing have driven the development of spoken language models (SLMs), designed to achieve universal audio understanding by jointly learning text and audio representations for a wide range of tasks.…

计算与语言 · 计算机科学 2025-10-31 Pedro Corrêa , João Lima , Victor Moreno , Lucas Ueda , Paula Dornhofer Paro Costa

The Kalman filter is a fundamental filtering algorithm that fuses noisy sensory data, a previous state estimate, and a dynamics model to produce a principled estimate of the current state. It assumes, and is optimal for, linear models and…

神经与进化计算 · 计算机科学 2021-04-30 Beren Millidge , Alexander Tschantz , Anil Seth , Christopher Buckley

Continuous dimensional speech emotion recognition captures affective variation along valence, arousal, and dominance, providing finer-grained representations than categorical approaches. Yet most multimodal methods rely solely on global…

声音 · 计算机科学 2026-01-27 Haoxun Li , Yuqing Sun , Hanlei Shi , Yu Liu , Leyuan Qu , Taihao Li

Although automatic emotion recognition (AER) has recently drawn significant research interest, most current AER studies use manually segmented utterances, which are usually unavailable for dialogue systems. This paper proposes integrating…

音频与语音处理 · 电气工程与系统科学 2023-08-15 Wen Wu , Chao Zhang , Philip C. Woodland

Speech Emotion Recognition (SER) is fundamental to affective computing and human-computer interaction, yet existing models struggle to generalize across diverse acoustic conditions. While Contrastive Language-Audio Pretraining (CLAP)…

声音 · 计算机科学 2025-07-08 Jiacheng Shi , Yanfu Zhang , Ye Gao

Recognizing emotional signals in speech has a significant impact on enhancing the effectiveness of human-computer interaction (HCI). This study introduces EmoAugNet, a hybrid deep learning framework, that incorporates Long Short-Term Memory…

声音 · 计算机科学 2025-08-11 Durjoy Chandra Paul , Gaurob Saha , Md Amjad Hossain

Speech Self-Supervised Learning (SSL) has demonstrated considerable efficacy in various downstream tasks. Nevertheless, prevailing self-supervised models often overlook the incorporation of emotion-related prior information, thereby…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Rui Liu , Zening Ma

Identifying emotion from speech is a non-trivial task pertaining to the ambiguous definition of emotion itself. In this work, we adopt a feature-engineering based approach to tackle the task of speech emotion recognition. Formalizing our…

机器学习 · 计算机科学 2019-04-15 Gaurav Sahu

Previous work on emotion recognition demonstrated a synergistic effect of combining several modalities such as auditory, visual, and transcribed text to estimate the affective state of a speaker. Among these, the linguistic modality is…

计算与语言 · 计算机科学 2019-03-01 Egor Lakomkin , Mohammad Ali Zamani , Cornelius Weber , Sven Magg , Stefan Wermter

Attention has become one of the most commonly used mechanisms in deep learning approaches. The attention mechanism can help the system focus more on the feature space's critical regions. For example, high amplitude regions can play an…

声音 · 计算机科学 2022-08-24 Junghun Kim , Yoojin An , Jihie Kim

In this paper, we propose to utilise diffusion models for data augmentation in speech emotion recognition (SER). In particular, we present an effective approach to utilise improved denoising diffusion probabilistic models (IDDPM) to…

声音 · 计算机科学 2023-05-22 Ibrahim Malik , Siddique Latif , Raja Jurdak , Björn Schuller

In this paper, we propose a multimodal framework for speech emotion recognition that leverages entropy-aware score selection to combine speech and textual predictions. The proposed method integrates a primary pipeline that consists of an…

声音 · 计算机科学 2025-08-29 ChenYi Chua , JunKai Wong , Chengxin Chen , Xiaoxiao Miao