English
Related papers

Related papers: CO-VADA: A Confidence-Oriented Voice Augmentation …

200 papers

Recent autoregressive transformer-based speech enhancement (SE) methods have shown promising results by leveraging advanced semantic understanding and contextual modeling of speech. However, these approaches often rely on complex…

Sound · Computer Science 2025-10-03 Luca A. Lanzendörfer , Frédéric Berdoz , Antonis Asonitis , Roger Wattenhofer

Medical diagnosis assistant (MDA) aims to build an interactive diagnostic agent to sequentially inquire about symptoms for discriminating diseases. However, since the dialogue records used to build a patient simulator are collected…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Junfan Lin , Keze Wang , Ziliang Chen , Xiaodan Liang , Liang Lin

Speech emotion recognition is an important component of any human centered system. But speech characteristics produced and perceived by a person can be influenced by a multitude of reasons, both desirable such as emotion, and undesirable…

Sound · Computer Science 2023-09-04 Mimansa Jaiswal , Emily Mower Provost

Diffusion-based generative models have recently gained attention in speech enhancement (SE), providing an alternative to conventional supervised methods. These models transform clean speech training samples into Gaussian noise centered at…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Jean-Eudes Ayilo , Mostafa Sadeghi , Romain Serizel

Speech Quality Assessment (SQA) and Continuous Speech Emotion Recognition (CSER) are two key tasks in speech technology, both relying on listener ratings. However, these ratings are inherently biased due to individual listener factors.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-22 Cheng-Hung Hu , Yusuke Yasuda , Akifumi Yoshimoto , Tomoki Toda

Imitation learning in robotics faces significant challenges in generalization due to the complexity of robotic environments and the high cost of data collection. We introduce RoCoDA, a novel method that unifies the concepts of invariance,…

Robotics · Computer Science 2025-05-21 Ezra Ameperosa , Jeremy A. Collins , Mrinal Jain , Animesh Garg

Keyword spotting and in particular Wake-Up-Word (WUW) detection is a very important task for voice assistants. A very common issue of voice assistants is that they get easily activated by background noise like music, TV or background speech…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-01 David Bonet , Guillermo Cámbara , Fernando López , Pablo Gómez , Carlos Segura , Jordi Luque

Speech emotion recognition (SER) systems often struggle in real-world environments, where ambient noise severely degrades their performance. This paper explores a novel approach that exploits prior knowledge of testing environments to…

Sound · Computer Science 2025-11-11 Seong-Gyun Leem , Daniel Fulford , Jukka-Pekka Onnela , David Gard , Carlos Busso

Personalized Voice Activity Detection (PVAD) systems activate only in response to a specific target speaker. Speaker-conditioning methods are employed to inject information about the target speaker into a VAD pipeline, to achieve…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-12 Mahsa Ghazvini Nejad , Hamed Jafarzadeh Asl , Amin Edraki , Mohammadreza Sadeghi , Masoud Asgharian , Yuanhao Yu , Vahid Partovi Nia

Deep learning models often learn and exploit spurious correlations in training data, using these non-target features to inform their predictions. Such reliance leads to performance degradation and poor generalization on unseen data. To…

Computation and Language · Computer Science 2025-11-21 Kyohoon Jin , Juhwan Choi , Jungmin Yun , Junho Lee , Soojin Jang , Youngbin Kim

In real-world environments, background noise significantly degrades the intelligibility and clarity of human speech. Audio-visual speech enhancement (AVSE) attempts to restore speech quality, but existing methods often fall short,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-27 Tassadaq Hussain , Kia Dashtipour , Yu Tsao , Amir Hussain

Speech emotion recognition (SER) has been one of the significant tasks in Human-Computer Interaction (HCI) applications. However, it is hard to choose the optimal features and deal with imbalance labeled data. In this article, we…

Sound · Computer Science 2021-09-21 Nhat Truong Pham , Duc Ngoc Minh Dang , Sy Dzung Nguyen

Speech Emotion Recognition (SER) systems often assume congruence between vocal emotion and lexical semantics. However, in real-world interactions, acoustic-semantic conflict is common yet overlooked, where the emotion conveyed by tone…

Sound · Computer Science 2026-01-09 Dawei Huang , Yongjie Lv , Ruijie Xiong , Chunxiang Jin , Xiaojiang Peng

Speech Emotion Recognition (SER) presents a significant yet persistent challenge in human-computer interaction. While deep learning has advanced spoken language processing, achieving high performance on limited datasets remains a critical…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Tai Vu

We propose EmoDistill, a novel speech emotion recognition (SER) framework that leverages cross-modal knowledge distillation during training to learn strong linguistic and prosodic representations of emotion from speech. During inference,…

Computation and Language · Computer Science 2024-03-18 Debaditya Shome , Ali Etemad

Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. Augmenting these systems with visual signals has the potential to improve robustness to noise. However, audio-visual…

Sound · Computer Science 2024-08-13 HyoJung Han , Mohamed Anwar , Juan Pino , Wei-Ning Hsu , Marine Carpuat , Bowen Shi , Changhan Wang

Speech deepfake detection (SDD) systems perform well on standard benchmarks datasets but often fail to generalize to expressive and emotional spoofing attacks. Many methods rely on spoof-heavy training data, learning dataset-specific…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-16 Aurosweta Mahapatra , Ismail Rasim Ulgen , Kong Aik Lee , Nicholas Andrews , Berrak Sisman

The human voice conveys unique characteristics of an individual, making voice biometrics a key technology for verifying identities in various industries. Despite the impressive progress of speaker recognition systems in terms of accuracy, a…

Sound · Computer Science 2022-08-24 Gianni Fenu , Giacomo Medda , Mirko Marras , Giacomo Meloni

In Speech Emotion Recognition (SER), emotional characteristics often appear in diverse forms of energy patterns in spectrograms. Typical attention neural network classifiers of SER are usually optimized on a fixed attention granularity. In…

Sound · Computer Science 2021-02-04 Mingke Xu , Fan Zhang , Xiaodong Cui , Wei Zhang

This study focuses on the First VoicePrivacy Attacker Challenge within the ICASSP 2025 Signal Processing Grand Challenge, which aims to develop speaker verification systems capable of determining whether two anonymized speech signals are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Yanzhe Zhang , Zhonghao Bi , Feiyang Xiao , Xuefeng Yang , Qiaoxi Zhu , Jian Guan