中文
相关论文

相关论文: Explainable Transformer-CNN Fusion for Noise-Robus…

200 篇论文

Speech is the most natural way of expressing ourselves as humans. Identifying emotion from speech is a nontrivial task due to the ambiguous definition of emotion itself. Speaker Emotion Recognition (SER) is essential for understanding human…

声音 · 计算机科学 2024-11-07 Pourya Jafarzadeh , Amir Mohammad Rostami , Padideh Choobdar

Speech-based machine learning systems are sensitive to noise, complicating reliable deployment in emotion recognition and voice pathology detection. We evaluate the robustness of a hybrid quantum machine learning model, quanvolutional…

声音 · 计算机科学 2026-01-07 Ha Tran , Bipasha Kashyap , Pubudu N. Pathirana

Speech emotion recognition (SER) often experiences reduced performance due to background noise. In addition, making a prediction on signals with only background noise could undermine user trust in the system. In this study, we propose a…

音频与语音处理 · 电气工程与系统科学 2023-09-06 Yu-Wen Chen , Julia Hirschberg , Yu Tsao

This paper presents a hybrid model combining Transformer and CNN for predicting the current waveform in signal lines. Unlike traditional approaches such as current source models, driver linear representations, waveform functional fitting,…

信号处理 · 电气工程与系统科学 2025-06-11 Junlang Huang , Hao Chen , Li Luo , Yong Cai , Lexin Zhang , Tianhao Ma , Yitian Zhang , Zhong Guan

The understanding of the surrounding environment plays a critical role in autonomous robotic systems, such as self-driving cars. Extensive research has been carried out concerning visual perception. Yet, to obtain a more complete perception…

音频与语音处理 · 电气工程与系统科学 2021-01-13 Karim Guirguis , Christoph Schorn , Andre Guntoro , Sherif Abdulatif , Bin Yang

This paper describes a submission to the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2) 2026, which addresses component-level deepfake detection using the CompSpoofV2 dataset, where speech and environmental sounds…

声音 · 计算机科学 2026-05-06 Khalid Zaman , Qixuan Huang , Muhammad Uzair , Masashi Unoki

Speech emotion recognition (SER) in naturalistic conditions presents a significant challenge for the speech processing community. Challenges include disagreement in labeling among annotators and imbalanced data distributions. This paper…

机器学习 · 计算机科学 2025-06-13 Thanathai Lertpetchpun , Tiantian Feng , Dani Byrd , Shrikanth Narayanan

Speech emotion recognition (SER) has been a popular research topic in human-computer interaction (HCI). As edge devices are rapidly springing up, applying SER to edge devices is promising for a huge number of HCI applications. Although deep…

声音 · 计算机科学 2023-05-12 Yi Chang , Zhao Ren , Thanh Tam Nguyen , Kun Qian , Björn W. Schuller

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. To improve robustness of speaker recognition system performance in…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Yanpei Shi , Qiang Huang , Thomas Hain

In recent years, speech enhancement (SE) has achieved impressive progress with the success of deep neural networks (DNNs). However, the DNN approach usually fails to generalize well to unseen environmental noise that is not included in the…

音频与语音处理 · 电气工程与系统科学 2020-04-09 Haoyu Li , Junichi Yamagishi

Accurate recognition of Emergency Vehicle (EV) sirens is critical for the integration of intelligent transportation systems, smart city monitoring systems, and autonomous driving technologies. Modern automatic solutions are limited by the…

声音 · 计算机科学 2025-07-01 Stefano Giacomelli , Marco Giordano , Claudia Rinaldi , Fabio Graziosi

Speech Emotion Recognition (SER) plays a pivotal role in enhancing human-computer interaction by enabling a deeper understanding of emotional states across a wide range of applications, contributing to more empathetic and effective…

音频与语音处理 · 电气工程与系统科学 2023-09-25 Amirali Soltani Tehrani , Niloufar Faridani , Ramin Toosi

Convolutional Neural Networks (CNNs) are effective models for reducing spectral variations and modeling spectral correlations in acoustic features for automatic speech recognition (ASR). Hybrid speech recognition systems incorporating CNNs…

This paper proposes a Convolutional Neural Network (CNN) inspired by Multitask Learning (MTL) and based on speech features trained under the joint supervision of softmax loss and center loss, a powerful metric learning strategy, for the…

声音 · 计算机科学 2019-09-04 Suraj Tripathi , Abhiram Ramesh , Abhay Kumar , Chirag Singh , Promod Yenigalla

Human emotion recognition plays an important role in human-computer interaction. In this paper, we present our approach to the Valence-Arousal (VA) Estimation Challenge, Expression (Expr) Classification Challenge, and Action Unit (AU)…

计算机视觉与模式识别 · 计算机科学 2023-09-07 Weiwei Zhou , Jiada Lu , Zhaolong Xiong , Weifeng Wang

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application…

声音 · 计算机科学 2024-12-18 Jinyi Mi , Xiaohan Shi , Ding Ma , Jiajun He , Takuya Fujimura , Tomoki Toda

Speech Emotion Recognition (SER) has become a growing focus of research in human-computer interaction. Spatiotemporal features play a crucial role in SER, yet current research lacks comprehensive spatiotemporal feature learning. This paper…

声音 · 计算机科学 2023-12-29 Mengbo Li , Yuanzhong Zheng , Dichucheng Li , Yulun Wu , Yaoxuan Wang , Haojun Fei

Voice disorders negatively impact the quality of daily life in various ways. However, accurately recognizing the category of pathological features from raw audio remains a considerable challenge due to the limited dataset. A promising…

声音 · 计算机科学 2024-10-08 Lipeng Shen , Yifan Xiong , Dongyue Guo , Wei Mo , Lingyu Yu , Hui Yang , Yi Lin

In this paper, we propose an innovative approach to perform speaker recognition by fusing two recently introduced deep neural networks (DNNs) namely - SincNet and X-Vector. The idea behind using SincNet filters on the raw speech waveform is…

计算与语言 · 计算机科学 2020-04-07 Mayank Tripathi , Divyanshu Singh , Seba Susan

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additive noise. The…

计算与语言 · 计算机科学 2015-03-31 Matthew Ager , Zoran Cvetkovic , Peter Sollich