中文
相关论文

相关论文: Non-linear frequency warping using constant-Q tran…

200 篇论文

Multimodal emotion recognition (MER) aims to infer human affect by jointly modeling audio and visual cues; however, existing approaches often struggle with temporal misalignment, weakly discriminative feature representations, and suboptimal…

多媒体 · 计算机科学 2026-01-21 Joe Dhanith P R , Shravan Venkatraman , Vigya Sharma , Santhosh Malarvannan

In this paper the task of emotion recognition from speech is considered. Proposed approach uses deep recurrent neural network trained on a sequence of acoustic features calculated over small speech intervals. At the same time special…

计算与语言 · 计算机科学 2018-07-06 Vladimir Chernykh , Pavel Prikhodko

In this work, we introduce SeQuiFi, a novel approach for mitigating catastrophic forgetting (CF) in speech emotion recognition (SER). SeQuiFi adopts a sequential class-finetuning strategy, where the model is fine-tuned incrementally on one…

音频与语音处理 · 电气工程与系统科学 2024-10-17 Sarthak Jain , Orchid Chetia Phukan , Swarup Ranjan Behera , Arun Balaji Buduru , Rajesh Sharma

Deriving an effective facial expression recognition component is important for a successful human-computer interaction system. Nonetheless, recognizing facial expression remains a challenging task. This paper describes a novel approach…

计算机视觉与模式识别 · 计算机科学 2019-04-15 Mundher Al-Shabi , Wooi Ping Cheah , Tee Connie

Connectionist Temporal Classification has recently attracted a lot of interest as it offers an elegant approach to building acoustic models (AMs) for speech recognition. The CTC loss function maps an input sequence of observable feature…

计算与语言 · 计算机科学 2017-08-16 Thomas Zenkel , Ramon Sanabria , Florian Metze , Jan Niehues , Matthias Sperber , Sebastian Stüker , Alex Waibel

Many neural speech enhancement and source separation systems operate in the time-frequency domain. Such models often benefit from making their Short-Time Fourier Transform (STFT) front-ends trainable. In current literature, these are…

音频与语音处理 · 电气工程与系统科学 2020-02-24 Jonah Casebeer , Umut Isik , Shrikant Venkataramani , Arvindh Krishnaswamy

Robustness against temporal variations is important for emotion recognition from speech audio, since emotion is ex-pressed through complex spectral patterns that can exhibit significant local dilation and compression on the time axis…

音频与语音处理 · 电气工程与系统科学 2020-03-10 Eric Guizzo , Tillman Weyde , Jack Barnett Leveson

This study proposes a novel approach for real-time facial expression recognition utilizing short-range Frequency-Modulated Continuous-Wave (FMCW) radar equipped with one transmit (Tx), and three receive (Rx) antennas. The system leverages…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Sabri Mustafa Kahya , Muhammet Sami Yavuz , Eckehard Steinbach

In this report we describe an ongoing line of research for solving single-channel source separation problems. Many monaural signal decomposition techniques proposed in the literature operate on a feature space consisting of a time-frequency…

声音 · 计算机科学 2015-04-29 Pablo Sprechmann , Joan Bruna , Yann LeCun

Speech emotion recognition (SER) classifies human emotions in speech with a computer model. Recently, performance in SER has steadily increased as deep learning techniques have adapted. However, unlike many domains that use speech data,…

声音 · 计算机科学 2024-09-09 Byunggun Kim , Younghun Kwon

Speech Emotion Recognition (SER) has emerged as a critical component of the next generation human-machine interfacing technologies. In this work, we propose a new dual-level model that predicts emotions based on both MFCC features and…

音频与语音处理 · 电气工程与系统科学 2020-07-24 Jianyou Wang , Michael Xue , Ryan Culhane , Enmao Diao , Jie Ding , Vahid Tarokh

We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that require large model capacity and substantial memory…

声音 · 计算机科学 2025-03-24 Tao Feng , Zhiyuan Zhao , Yifan Xie , Yuqi Ye , Xiangyang Luo , Xun Guan , Yu Li

Emotion recognition using Electroencephalogram (EEG) signals has emerged as a significant research challenge in affective computing and intelligent interaction. However, effectively combining global and local features of EEG signals to…

信号处理 · 电气工程与系统科学 2023-05-10 Wei Lu , Hua Ma , Tien-Ping Tan

Detecting emotions directly from a speech signal plays an important role in effective human-computer interactions. Existing speech emotion recognition models require massive computational and storage resources, making them hard to implement…

音频与语音处理 · 电气工程与系统科学 2021-10-08 Arya Aftab , Alireza Morsali , Shahrokh Ghaemmaghami , Benoit Champagne

Speech Emotion Recognition (SER) plays a pivotal role in enhancing human-computer interaction by enabling a deeper understanding of emotional states across a wide range of applications, contributing to more empathetic and effective…

音频与语音处理 · 电气工程与系统科学 2023-09-25 Amirali Soltani Tehrani , Niloufar Faridani , Ramin Toosi

Over the past few years, speech enhancement methods based on deep learning have greatly surpassed traditional methods based on spectral subtraction and spectral estimation. Many of these new techniques operate directly in the the short-time…

音频与语音处理 · 电气工程与系统科学 2020-08-28 Jean-Marc Valin , Umut Isik , Neerad Phansalkar , Ritwik Giri , Karim Helwani , Arvindh Krishnaswamy

We introduce the Latent Fourier Transform (LatentFT), a framework that provides novel frequency-domain controls for generative music models. LatentFT combines a diffusion autoencoder with a latent-space Fourier transform to separate musical…

声音 · 计算机科学 2026-04-21 Mason Wang , Cheng-Zhi Anna Huang

For supernovae gravitational wave signal analysis which intend to reconstruct supernova gravitational waves waveforms, we compare the performance of short-time Fourier transform (STFT), the synchroextracting transform (SET) and…

高能天体物理现象 · 物理学 2020-10-28 Zhuotao Li , Xilong Fan , Gang Yu

We propose a novel transfer learning method for speech emotion recognition allowing us to obtain promising results when only few training data is available. With as low as 125 examples per emotion class, we were able to reach a higher…

机器学习 · 计算机科学 2020-11-12 Jonathan Boigne , Biman Liyanage , Ted Östrem

Purpose: To develop a robust deep learning framework for non-contrast-enhanced functional lung MRI, overcoming the limitations of spectral decomposition in the presence of physiological non-stationarity. Methods: We introduce VQ-Wave…

医学物理 · 物理学 2026-05-05 Grzegorz Bauman , Pavlos Panos , Philipp Latzin , Oliver Bieri