English
Related papers

Related papers: Prosody-Adaptable Audio Codecs for Zero-Shot Voice…

200 papers

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces the resulting…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Yang Yang , Yury Kartynnik , Yunpeng Li , Jiuqiang Tang , Xing Li , George Sung , Matthias Grundmann

This paper introduces Taco-VC, a novel architecture for voice conversion based on Tacotron synthesizer, which is a sequence-to-sequence with attention model. The training of multi-speaker voice conversion systems requires a large number of…

Sound · Computer Science 2020-06-22 Roee Levy Leshem , Raja Giryes

Zero-shot voice conversion aims to transform a source speech utterance to match the timbre of a reference speech from an unseen speaker. Traditional approaches struggle with timbre leakage, insufficient timbre representation, and mismatches…

Sound · Computer Science 2024-11-18 Songting Liu

This paper presents a new voice conversion model capable of transforming both speaking and singing voices. It addresses key challenges in current systems, such as conveying emotions, managing pronunciation and accent changes, and…

Sound · Computer Science 2024-12-12 Sowmya Cheripally

Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to reduce the amount of unexplained variation in training data…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Devang S Ram Mohan , Vivian Hu , Tian Huey Teh , Alexandra Torresquintero , Christopher G. R. Wallis , Marlene Staib , Lorenzo Foglianti , Jiameng Gao , Simon King

Emotional voice conversion (EVC) aims to modify the emotional style of speech while preserving its linguistic content. In practical EVC, controllability, the ability to independently control speaker identity and emotional style using…

Sound · Computer Science 2025-08-12 Jinsung Yoon , Wooyeol Jeong , Jio Gim , Young-Joo Suh

Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Recently, AutoVC, a conditional autoencoder based method, achieved excellent conversion results by disentangling the speaker identity…

Sound · Computer Science 2022-08-09 Huaizhen Tang , Xulong Zhang , Jianzong Wang , Ning Cheng , Zhen Zeng , Edward Xiao , Jing Xiao

Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling diverse audio-text…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-27 Xiaofei Wang , Manthan Thakker , Zhuo Chen , Naoyuki Kanda , Sefik Emre Eskimez , Sanyuan Chen , Min Tang , Shujie Liu , Jinyu Li , Takuya Yoshioka

Voice conversion technologies have been greatly improved in recent years with the help of deep learning, but their capabilities of producing natural sounding utterances in different conditions remain unclear. In this paper, we gave a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-04 Tzu-hsien Huang , Jheng-hao Lin , Chien-yu Huang , Hung-yi Lee

This paper proposes a novel voice conversion (VC) method based on non-autoregressive sequence-to-sequence (NAR-S2S) models. Inspired by the great success of NAR-S2S models such as FastSpeech in text-to-speech (TTS), we extend the…

Sound · Computer Science 2021-04-15 Tomoki Hayashi , Wen-Chin Huang , Kazuhiro Kobayashi , Tomoki Toda

Voice conversion is an increasingly popular technology, and the growing number of real-time applications requires models with streaming conversion capabilities. Unlike typical (non-streaming) voice conversion, which can leverage the entire…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-01 Ziqian Ning , Yuepeng Jiang , Pengcheng Zhu , Jixun Yao , Shuai Wang , Lei Xie , Mengxiao Bi

Expressive speech synthesis models are trained by adding corpora with diverse speakers, various emotions, and different speaking styles to the dataset, in order to control various characteristics of speech and generate the desired voice. In…

Sound · Computer Science 2023-07-21 Daegyeom Kim , Seongho Hong , Yong-Hoon Choi

Generating expressive audio performances from music scores requires models to capture both instrument acoustics and human interpretation. Traditional music performance synthesis pipelines follow a two-stage approach, first generating…

Sound · Computer Science 2025-07-14 Jingjing Tang , Xin Wang , Zhe Zhang , Junichi Yamagishi , Geraint Wiggins , George Fazekas

The prosodic aspects of speech signals produced by current text-to-speech systems are typically averaged over training material, and as such lack the variety and liveliness found in natural speech. To avoid monotony and averaged prosody…

Computation and Language · Computer Science 2019-06-05 Vincent Wan , Chun-an Chan , Tom Kenter , Jakub Vit , Rob Clark

This paper proposes a novel vision-integrated neural speech codec (VNSC), which aims to enhance speech coding quality by leveraging visual modality information. In VNSC, the image analysis-synthesis module extracts visual features from lip…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-30 Yao Guo , Yang Ai , Rui-Chen Zheng , Hui-Peng Du , Xiao-Hang Jiang , Zhen-Hua Ling

Unsupervised representation learning of speech has been of keen interest in recent years, which is for example evident in the wide interest of the ZeroSpeech challenges. This work presents a new method for learning frame level…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-18 Mingjie Chen , Thomas Hain

We introduce a novel sequence-to-sequence (seq2seq) voice conversion (VC) model based on the Transformer architecture with text-to-speech (TTS) pretraining. Seq2seq VC models are attractive owing to their ability to convert prosody. While…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-17 Wen-Chin Huang , Tomoki Hayashi , Yi-Chiao Wu , Hirokazu Kameoka , Tomoki Toda

Codec-based language models (LMs) have revolutionized text-to-speech (TTS). However, standard codecs entangle timbre and prosody, which hinders independent control in continuation-based LMs. To tackle this challenge, we propose…

Sound · Computer Science 2026-01-06 Tao Li , Wenshuo Ge , Zhichao Wang , Zihao Cui , Yong Ma , Yingying Gao , Chao Deng , Shilei Zhang , Junlan Feng

Vision-Language Models (VLMs) have demonstrated impressive multimodal capabilities in learning joint representations of visual and textual data, making them powerful tools for tasks such as Compositional Zero-Shot Learning (CZSL). CZSL…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Kyle Stein , Arash Mahyari , Guillermo Francia , Eman El-Sheikh

Emotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. The prior studies on emotional voice conversion are mostly carried out under the…

Sound · Computer Science 2020-10-14 Kun Zhou , Berrak Sisman , Mingyang Zhang , Haizhou Li