English
Related papers

Related papers: Lightweight and perceptually-guided voice conversi…

200 papers

Deep learning based speech denoising still suffers from the challenge of improving perceptual quality of enhanced signals. We introduce a generalized framework called Perceptual Ensemble Regularization Loss (PERL) built on the idea of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-23 Saurabh Kataria , Jesús Villalba , Najim Dehak

The evaluation of synthetic and processed speech has long been a cornerstone of audio engineering and speech science. Although subjective listening tests remain the gold standard for assessing perceptual quality and intelligibility, their…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-05 Yu Tsao

Sequence-to-sequence (seq2seq) voice conversion (VC) models have greater potential in converting electrolaryngeal (EL) speech to normal speech (EL2SP) compared to conventional VC models. However, EL2SP based on seq2seq VC requires a…

Sound · Computer Science 2022-10-20 Ding Ma , Lester Phillip Violeta , Kazuhiro Kobayashi , Tomoki Toda

We propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on factorizing speech into explicitly disentangled representations that…

Large Language Models (LLMs) are one of the most promising technologies for the next era of speech generation systems, due to their scalability and in-context learning capabilities. Nevertheless, they suffer from multiple stability issues…

Emotional state recognition through speech is being a very interesting research topic nowadays. Using subliminal information of speech, denominated as prosody, it is possible to recognize the emotional state of the person. One of the main…

Computer Vision and Pattern Recognition · Computer Science 2014-03-20 Inma Mohino-Herranz , Roberto Gil-Pita , Sagrario Alonso-Diaz , Manuel Rosa-Zurera

Expressive voice conversion aims to transfer both speaker identity and expressive attributes from a target speech to a given source speech. In this work, we improve over a self-supervised, non-autoregressive framework with a conditional…

Sound · Computer Science 2025-06-05 Seymanur Akti , Tuan Nam Nguyen , Alexander Waibel

This work adapts two recent architectures of generative models and evaluates their effectiveness for the conversion of whispered speech to normal speech. We incorporate the normal target speech into the training criterion of…

Emotion recognition from speech is a challenging task that requires capturing both linguistic and paralinguistic cues, with critical applications in human-computer interaction and mental health monitoring. Recent works have highlighted the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-21 Hugo Thimonier , Antony Perzo , Renaud Seguier

Realistic emotional voice conversion (EVC) aims to enhance emotional diversity of converted audios, making the synthesized voices more authentic and natural. To this end, we propose Emotional Intensity-aware Network (EINet), dynamically…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-23 Tianhua Qi , Shiyan Wang , Cheng Lu , Yan Zhao , Yuan Zong , Wenming Zheng

Unvoiced electromyography (EMG) is an effective communication tool for individuals unable to produce vocal speech. However, most prior methods rely on paired voiced and unvoiced EMG signals, along with speech data, for EMG-to-text…

Computation and Language · Computer Science 2025-06-03 Payal Mohapatra , Akash Pandey , Xiaoyuan Zhang , Qi Zhu

Prosody modeling is important, but still challenging in expressive voice conversion. As prosody is difficult to model, and other factors, e.g., speaker, environment and content, which are entangled with prosody in speech, should be removed…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-04 Wendong Gan , Bolong Wen , Ying Yan , Haitao Chen , Zhichao Wang , Hongqiang Du , Lei Xie , Kaixuan Guo , Hai Li

Emotional voice conversion (EVC) seeks to convert the emotional state of an utterance while preserving the linguistic content and speaker identity. In EVC, emotions are usually treated as discrete categories overlooking the fact that speech…

Sound · Computer Science 2022-07-19 Kun Zhou , Berrak Sisman , Rajib Rana , Björn W. Schuller , Haizhou Li

Speech enhancement (SE) aims to improve the clarity, intelligibility, and quality of speech signals for various speech enabled applications. However, air-conducted (AC) speech is highly susceptible to ambient noise, particularly in low…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Fuyuan Feng , Longting Xu , Rohan Kumar Das

Recent studies have shown that post-deployment adaptation can improve the robustness of speech enhancement models in unseen noise conditions. However, existing methods often incur prohibitive computational and memory costs, limiting their…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Longbiao Cheng , Shih-Chii Liu

In audio signal processing, learnable front-ends have shown strong performance across diverse tasks by optimizing task-specific representation. However, their parameters remain fixed once trained, lacking flexibility during inference and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-29 Hanyu Meng , Vidhyasaharan Sethu , Eliathamby Ambikairajah , Qiquan Zhang , Haizhou Li

Variational autoencoder-based voice conversion (VAE-VC) has the advantage of requiring only pairs of speeches and speaker labels for training. Unlike the majority of the research in VAE-VC which focuses on utilizing auxiliary losses or…

Sound · Computer Science 2021-12-07 Kei Akuzawa , Kotaro Onishi , Keisuke Takiguchi , Kohki Mametani , Koichiro Mori

Emotional voice conversion (EVC) involves modifying various acoustic characteristics, such as pitch and spectral envelope, to match a desired emotional state while preserving the speaker's identity. Existing EVC methods often rely on text…

Sound · Computer Science 2025-01-22 Hyung-Seok Oh , Sang-Hoon Lee , Deok-Hyeon Cho , Seong-Whan Lee

Attention-based contextual biasing approaches have shown significant improvements in the recognition of generic and/or personal rare-words in End-to-End Automatic Speech Recognition (E2E ASR) systems like neural transducers. These…

Computation and Language · Computer Science 2023-05-10 Xuandi Fu , Kanthashree Mysore Sathyendra , Ankur Gandhe , Jing Liu , Grant P. Strimel , Ross McGowan , Athanasios Mouchtaris

In recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such representations are still vulnerable to emotion variability. To…

Sound · Computer Science 2025-05-27 Jingguang Tian , Xinhui Hu , Xinkang Xu