English
Related papers

Related papers: HiFi-VC: High Quality ASR-Based Voice Conversion

200 papers

We propose noise-robust voice conversion (VC) which takes into account the recording quality and environment of noisy source speech. Conventional denoising training improves the noise robustness of a VC model by learning noisy-to-clean VC…

Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study…

Sound · Computer Science 2026-01-09 Prajwal Chinchmalatpure , Suyash Chinchmalatpure , Siddharth Chavan

In this paper we propose a new cross-lingual Voice Conversion (VC) approach which can generate all speech parameters (MCEP, LF0, BAP) from one DNN model using PPGs (Phonetic PosteriorGrams) extracted from inputted speech using several ASR…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-29 Qinghua Sun , Kenji Nagamatsu

Zero-shot voice conversion (VC) aims to transfer timbre from a source speaker to any unseen target speaker while preserving linguistic content. Growing application scenarios demand models with streaming inference capabilities. This has…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-23 Guobin Ma , Jixun Yao , Ziqian Ning , Yuepeng Jiang , Lingxin Xiong , Lei Xie , Pengcheng Zhu

This paper presents a method for end-to-end cross-lingual text-to-speech (TTS) which aims to preserve the target language's pronunciation regardless of the original speaker's language. The model used is based on a non-attentive Tacotron…

The performance of child speech recognition is generally less satisfactory compared to adult speech due to limited amount of training data. Significant performance degradation is expected when applying an automatic speech recognition (ASR)…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-26 Wei Liu , Jingyu Li , Tan Lee

Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-09 Xinyu Wang , Haotian Jiang , Haolin Huang , Yu Fang , Mengjie Xu , Qian Wang

We introduce a novel sequence-to-sequence (seq2seq) voice conversion (VC) model based on the Transformer architecture with text-to-speech (TTS) pretraining. Seq2seq VC models are attractive owing to their ability to convert prosody. While…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-17 Wen-Chin Huang , Tomoki Hayashi , Yi-Chiao Wu , Hirokazu Kameoka , Tomoki Toda

Visual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visually similar lip…

Artificial Intelligence · Computer Science 2024-06-19 Young Jin Ahn , Jungwoo Park , Sangha Park , Jonghyun Choi , Kee-Eung Kim

Typically, singing voice conversion (SVC) depends on an embedding vector, extracted from either a speaker lookup table (LUT) or a speaker recognition network (SRN), to model speaker identity. However, singing contains more expressive…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-07 Xu Li , Shansong Liu , Ying Shan

Privacy-preserving voice conversion aims to remove only the attributes of speech audio that convey identity information, keeping other speech characteristics intact. This paper presents a mechanism for privacy-preserving voice conversion…

Sound · Computer Science 2024-09-24 Jacob J Webber , Oliver Watts , Gustav Eje Henter , Jennifer Williams , Simon King

We propose TES-VC (Text-driven Environment and Speaker controllable Voice Conversion), a text-driven voice conversion framework with independent control of speaker timbre and environmental acoustics. TES-VC processes simultaneous text…

Sound · Computer Science 2025-06-16 Jiawei Jin , Zhihan Yang , Yixuan Zhou , Zhiyong Wu

In this study, we try to address the problem of leveraging visual signals to improve Automatic Speech Recognition (ASR), also known as visual context-aware ASR (VC-ASR). We explore novel VC-ASR approaches to leverage video and text…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-10 Shahram Ghorbani , Yashesh Gaur , Yu Shi , Jinyu Li

Automatic speech recognition (ASR) systems become increasingly efficient thanks to new advances in neural network training like self-supervised learning. However, they are known to be unfair toward certain groups, for instance, people…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-07 Lucas Maison , Yannick Estève

Voice conversion models modify timbre while preserving paralinguistic features, enabling applications like dubbing and identity protection. However, most VC systems require access to target utterances, limiting their use when target data is…

Sound · Computer Science 2025-11-11 Meiying Melissa Chen , Zhenyu Wang , Zhiyao Duan

Humans are adept at leveraging visual cues from lip movements for recognizing speech in adverse listening conditions. Audio-Visual Speech Recognition (AVSR) models follow similar approach to achieve robust speech recognition in noisy…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-24 Maxime Burchi , Krishna C. Puvvada , Jagadeesh Balam , Boris Ginsburg , Radu Timofte

The rising trend of using voice as a means of interacting with smart devices has sparked worries over the protection of users' privacy and data security. These concerns have become more pressing, especially after the European Union's…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-19 Suhita Ghosh , Yamini Sinha , Ingo Siegert , Sebastian Stober

Conversational automatic speech recognition (ASR) is a task to recognize conversational speech including multiple speakers. Unlike sentence-level ASR, conversational ASR can naturally take advantages from specific characteristics of…

Sound · Computer Science 2022-02-18 Kun Wei , Yike Zhang , Sining Sun , Lei Xie , Long Ma

This paper proposes RefXVC, a method for cross-lingual voice conversion (XVC) that leverages reference information to improve conversion performance. Previous XVC works generally take an average speaker embedding to condition the speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-25 Mingyang Zhang , Yi Zhou , Yi Ren , Chen Zhang , Xiang Yin , Haizhou Li

An audiovisual speaker conversion method is presented for simultaneously transforming the facial expressions and voice of a source speaker into those of a target speaker. Transforming the facial and acoustic features together makes it…

Audio and Speech Processing · Electrical Eng. & Systems 2018-12-04 Fuming Fang , Xin Wang , Junichi Yamagishi , Isao Echizen