English
Related papers

Related papers: MSM-VC: High-fidelity Source Style Transfer for No…

200 papers

We introduce DISSC, a novel, lightweight method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner. Unlike DISSC, most voice conversion (VC) methods focus primarily on timbre, and…

Sound · Computer Science 2023-10-20 Gallil Maimon , Yossi Adi

In this paper, a neural network named Sequence-to-sequence ConvErsion NeTwork (SCENT) is presented for acoustic modeling in voice conversion. At training stage, a SCENT model is estimated by aligning the feature sequences of source and…

Sound · Computer Science 2020-01-14 Jing-Xuan Zhang , Zhen-Hua Ling , Li-Juan Liu , Yuan Jiang , Li-Rong Dai

Recent Speech Large Language Models~(LLMs) have achieved impressive capabilities in end-to-end speech interaction. However, the prevailing autoregressive paradigm imposes strict serial constraints, limiting generation efficiency and…

Computation and Language · Computer Science 2026-02-10 Ziyang Cheng , Yuhao Wang , Heyang Liu , Ronghua Wu , Qunshan Gu , Yanfeng Wang , Yu Wang

We introduce HybridVC, a voice conversion (VC) framework built upon a pre-trained conditional variational autoencoder (CVAE) that combines the strengths of a latent model with contrastive learning. HybridVC supports text and audio prompts,…

Sound · Computer Science 2024-09-26 Xinlei Niu , Jing Zhang , Charles Patrick Martin

This paper aims to synthesize the target speaker's speech with desired speaking style and emotion by transferring the style and emotion from reference speech recorded by other speakers. We address this challenging problem with a two-stage…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-15 Xinfa Zhu , Yi Lei , Kun Song , Yongmao Zhang , Tao Li , Lei Xie

Voice conversion aims to convert source speech into a target voice using recordings of the target speaker as a reference. Newer models are producing increasingly realistic output. But what happens when models are fed with non-standard data,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-13 Matthew Baas , Herman Kamper

This paper introduces an area-based source separation method designed for virtual meeting scenarios. The aim is to preserve speech signals from an unspecified number of sources within a defined spatial area in front of a linear microphone…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-20 Martin Strauss , Okan Köpüklü

Recent advances in discrete audio codecs have significantly improved speech representation modeling, while codec language models have enabled in-context learning for zero-shot speech synthesis. Inspired by this, we propose a voice…

Sound · Computer Science 2025-09-30 Junchuan Zhao , Xintong Wang , Ye Wang

This paper presents a novel framework to build a voice conversion (VC) system by learning from a text-to-speech (TTS) synthesis system, that is called TTS-VC transfer learning. We first develop a multi-speaker speech synthesis system with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-07 Mingyang Zhang , Yi Zhou , Li Zhao , Haizhou Li

Large Language Models (LLMs) are one of the most promising technologies for the next era of speech generation systems, due to their scalability and in-context learning capabilities. Nevertheless, they suffer from multiple stability issues…

Singing Voice Conversion (SVC) has emerged as a significant subfield of Voice Conversion (VC), enabling the transformation of one singer's voice into another while preserving musical elements such as melody, rhythm, and timbre. Traditional…

Sound · Computer Science 2025-01-22 Yubo Huang , Xin Lai , Muyang Ye , Anran Zhu , Zixi Wang , Jingzehua Xu , Shuai Zhang , Zhiyuan Zhou , Weijie Niu

This paper explores predicting suitable prosodic features for fine-grained emotion analysis from the discourse-level text. To obtain fine-grained emotional prosodic features as predictive values for our model, we extract a phoneme-level…

Sound · Computer Science 2023-09-22 Xianhao Wei , Jia Jia , Xiang Li , Zhiyong Wu , Ziyi Wang

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…

Sound · Computer Science 2025-06-27 Kehan Sui , Jinxu Xiang , Fang Jin

Despite the success of style transfer in image processing, it has seen limited progress in natural language generation. Part of the problem is that content is not as easily decoupled from style in the text domain. Curiously, in the field of…

Computation and Language · Computer Science 2019-11-11 Katy Gero , Chris Kedzie , Jonathan Reeve , Lydia Chilton

In voice conversion (VC) applications, diffusion and flow-matching models have exhibited exceptional speech quality and speaker similarity performances. However, they are limited by slow conversion owing to their iterative inference.…

Sound · Computer Science 2026-02-23 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Yuto Kondo

One-shot voice conversion aims to change the timbre of any source speech to match that of the unseen target speaker with only one speech sample. Existing methods face difficulties in satisfactory speech representation disentanglement and…

Sound · Computer Science 2024-11-26 Pengcheng Li , Jianzong Wang , Xulong Zhang , Yong Zhang , Jing Xiao , Ning Cheng

Voice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio…

Sound · Computer Science 2024-01-19 Yimin Deng , Huaizhen Tang , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

Speech discrete representation has proven effective in various downstream applications due to its superior compression rate of the waveform, fast convergence during training, and compatibility with other modalities. Discrete units extracted…

Sound · Computer Science 2024-06-17 Jiatong Shi , Xutai Ma , Hirofumi Inaguma , Anna Sun , Shinji Watanabe

Singing voice conversion aims to transform a source singing voice into that of a target singer while preserving the original lyrics, melody, and various vocal techniques. In this paper, we propose a high-fidelity singing voice conversion…

Sound · Computer Science 2025-01-07 Yiquan Zhou , Wenyu Wang , Hongwu Ding , Jiacheng Xu , Jihua Zhu , Xin Gao , Shihao Li

Emotional voice conversion (VC) aims to convert a neutral voice to an emotional (e.g. happy) one while retaining the linguistic information and speaker identity. We note that the decoupling of emotional features from other speech…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-05 Zhaojie Luo , Shoufeng Lin , Rui Liu , Jun Baba , Yuichiro Yoshikawa , Ishiguro Hiroshi