English
Related papers

Related papers: SemAlignVC: Enhancing zero-shot timbre conversion …

200 papers

Protecting speaker identity is crucial for online voice applications, yet streaming speaker anonymization (SA) remains underexplored. Recent research has demonstrated that neural audio codec (NAC) provides superior speaker feature…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-06 Nikita Kuzmin , Songting Liu , Kong Aik Lee , Eng Siong Chng

Zero-shot multi-speaker TTS aims to synthesize speech with the voice of a chosen target speaker without any fine-tuning. Prevailing methods, however, encounter limitations at adapting to new speakers of out-of-domain settings, primarily due…

Sound · Computer Science 2024-03-06 Yejin Jeon , Yunsu Kim , Gary Geunbae Lee

Singing Voice Synthesis (SVS) aims to generate expressive vocal performances from structured musical inputs such as lyrics and pitch sequences. While recent progress in discrete codec-based speech synthesis has enabled zero-shot generation…

Sound · Computer Science 2026-04-14 Junchuan Zhao , Wei Zeng , Tianle Lyu , Ye Wang

Joint audio-video generation aims to synthesize synchronized multisensory content, yet current unified models struggle with fine-grained acoustic control, particularly for identity-preserving speech. Existing approaches either suffer from…

Sound · Computer Science 2026-01-09 Chunyu Qiang , Jun Wang , Xiaopeng Wang , Kang Yin , Yuxin Guo

Recent advancements highlight the potential of end-to-end real-time spoken dialogue systems, showcasing their low latency and high quality. In this paper, we introduce SLAM-Omni, a timbre-controllable, end-to-end voice interaction system…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-23 Wenxi Chen , Ziyang Ma , Ruiqi Yan , Yuzhe Liang , Xiquan Li , Ruiyang Xu , Zhikang Niu , Yanqiao Zhu , Yifan Yang , Zhanxun Liu , Kai Yu , Yuxuan Hu , Jinyu Li , Yan Lu , Shujie Liu , Xie Chen

Recently, there have been significant advancements in voice conversion, resulting in high-quality performance. However, there are still two critical challenges in this field. Firstly, current voice conversion methods have limited robustness…

Sound · Computer Science 2024-08-13 Le Xu , Jiangyan Yi , Tao Wang , Yong Ren , Rongxiu Zhong , Zhengqi Wen , Jianhua Tao

Text-based voice editing (TBVE) uses synthetic output from text-to-speech (TTS) systems to replace words in an original recording. Recent work has used neural models to produce edited speech that is similar to the original speech in terms…

Sound · Computer Science 2022-10-31 Jason Fong , Yun Wang , Prabhav Agrawal , Vimal Manohar , Jilong Wu , Thilo Köhler , Qing He

Traditional voice conversion methods rely on parallel recordings of multiple speakers pronouncing the same sentences. For real-world applications however, parallel data is rarely available. We propose MelGAN-VC, a voice conversion method…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-06 Marco Pasini

The rapid advancement of generative AI has made it increasingly challenging to distinguish between deepfake audio and authentic human speech. To overcome the limitations of passive detection methods, we propose StreamMark, a novel deep…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Zhentao Liu , Milos Cernak

Prior works have demonstrated zero-shot text-to-speech by using a generative language model on audio tokens obtained via a neural audio codec. It is still challenging, however, to adapt them to low-latency scenarios. In this paper, we…

Sound · Computer Science 2024-06-11 Trung Dang , David Aponte , Dung Tran , Kazuhito Koishida

The goal of voice conversion (VC) is to convert input voice to match the target speaker's voice while keeping text and prosody intact. VC is usually used in entertainment and speaking-aid systems, as well as applied for speech data…

Sound · Computer Science 2022-04-01 A. Kashkin , I. Karpukhin , S. Shishkin

We propose TES-VC (Text-driven Environment and Speaker controllable Voice Conversion), a text-driven voice conversion framework with independent control of speaker timbre and environmental acoustics. TES-VC processes simultaneous text…

Sound · Computer Science 2025-06-16 Jiawei Jin , Zhihan Yang , Yixuan Zhou , Zhiyong Wu

As a foundational technology for intelligent human-computer interaction, voice conversion (VC) seeks to transform speech from any source timbre into any target timbre. Traditional voice conversion methods based on Generative Adversarial…

Sound · Computer Science 2025-06-11 Wenhan Yao , Fen Xiao , Xiarun Chen , Jia Liu , YongQiang He , Weiping Wen

Learning to segment images purely by relying on the image-text alignment from web data can lead to sub-optimal performance due to noise in the data. The noise comes from the samples where the associated text does not correlate with the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-08 Yash Patel , Yusheng Xie , Yi Zhu , Srikar Appalaraju , R. Manmatha

The goal of accent conversion (AC) is to convert the accent of speech into the target accent while preserving the content and speaker identity. AC enables a variety of applications, such as language learning, speech content creation, and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-11 Dongya Jia , Qiao Tian , Kainan Peng , Jiaxin Li , Yuanzhe Chen , Mingbo Ma , Yuping Wang , Yuxuan Wang

This paper focuses on using voice conversion (VC) to improve the speech intelligibility of surgical patients who have had parts of their articulators removed. Due to the difficulty of data collection, VC without parallel data is highly…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-26 Li-Wei Chen , Hung-Yi Lee , Yu Tsao

General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting. Unlike supervised classification benchmarks that measure adaptability…

Sound · Computer Science 2025-12-12 Maris Basha , Anja Zai , Sabine Stoll , Richard Hahnloser

Recently, more and more zero-shot voice conversion algorithms have been proposed. As a fundamental part of zero-shot voice conversion, speaker embeddings are the key to improving the converted speech's speaker similarity. In this paper, we…

Sound · Computer Science 2022-03-21 Ruitong Xiao , Haitong Zhang , Yue Lin

Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, prevailing methods augment acoustic…

Sound · Computer Science 2026-01-28 Xin Zhang , Lin Li , Xiangni Lu , Jianquan Liu , Kong Aik Lee

Zero-shot multi-speaker text-to-speech (ZSM-TTS) models aim to generate a speech sample with the voice characteristic of an unseen speaker. The main challenge of ZSM-TTS is to increase the overall speaker similarity for unseen speakers. One…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-28 Byoung Jin Choi , Myeonghun Jeong , Joun Yeop Lee , Nam Soo Kim