English
Related papers

Related papers: SemAlignVC: Enhancing zero-shot timbre conversion …

200 papers

With recent advancements in voice cloning, the performance of speech synthesis for a target speaker has been rendered similar to the human level. However, autoregressive voice cloning systems still suffer from text alignment failures,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-27 Artem Gorodetskii , Ivan Ozhiganov

In this paper, we study zero-shot learning in audio classification via semantic embeddings extracted from textual labels and sentence descriptions of sound classes. Our goal is to obtain a classifier that is capable of recognizing audio…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-12 Huang Xie , Tuomas Virtanen

With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modalities. Recent speech tokenization approaches aim to isolate…

Computation and Language · Computer Science 2025-06-23 Daejin Jo , Jeeyoung Yun , Byungseok Roh , Sungwoong Kim

We introduce MaskVCT, a zero-shot voice conversion (VC) model that offers multi-factor controllability through multiple classifier-free guidances (CFGs). While previous VC models rely on a fixed conditioning scheme, MaskVCT integrates…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-12 Junhyeok Lee , Helin Wang , Yaohan Guan , Thomas Thebaud , Laureano Moro-Velazquez , Jesús Villalba , Najim Dehak

In recent years, large-scale pre-trained speech language models (SLMs) have demonstrated remarkable advancements in various generative speech modeling applications, such as text-to-speech synthesis, voice conversion, and speech enhancement.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-19 Yinghao Aaron Li , Cong Han , Nima Mesgarani

Singing Voice Conversion (SVC) aims to transform a source singing voice into a target singer while preserving lyrics and melody. Most existing SVC methods depend on F0 extractors to capture the lead melody from clean vocals. However, no…

Sound · Computer Science 2026-05-13 Chen Geng , Meng Chen , Ruohua Zhou , Ruolan Liu , Weifeng Zhao

Cross-lingual voice conversion (VC) is a task that aims to synthesize target voices with the same content while source and target speakers speak in different languages. Its challenge lies in the fact that the source and target data are…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-01 Che-Jui Chang

We observe that zero-shot appearance transfer with large-scale image generation models faces a significant challenge: Attention Leakage. This challenge arises when the semantic mapping between two images is captured by the Query-Key…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Namu Kim , Wonbin Kweon , Minsoo Kim , Hwanjo Yu

Typically, singing voice conversion (SVC) depends on an embedding vector, extracted from either a speaker lookup table (LUT) or a speaker recognition network (SRN), to model speaker identity. However, singing contains more expressive…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-07 Xu Li , Shansong Liu , Ying Shan

Voice conversion (VC) techniques aim to modify speaker identity of an utterance while preserving the underlying linguistic information. Most VC approaches ignore modeling of the speaking style (e.g. emotion and emphasis), which may contain…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-20 Songxiang Liu , Yuewen Cao , Shiyin Kang , Na Hu , Xunying Liu , Dan Su , Dong Yu , Helen Meng

We propose a parallel-data-free voice-conversion (VC) method that can learn a mapping from source to target speech without relying on parallel data. The proposed method is general purpose, high quality, and parallel-data free and works…

Machine Learning · Statistics 2017-12-21 Takuhiro Kaneko , Hirokazu Kameoka

Voice conversion models modify timbre while preserving paralinguistic features, enabling applications like dubbing and identity protection. However, most VC systems require access to target utterances, limiting their use when target data is…

Sound · Computer Science 2025-11-11 Meiying Melissa Chen , Zhenyu Wang , Zhiyao Duan

Recent work in cross-lingual semantic parsing has successfully applied machine translation to localize parsers to new languages. However, these advances assume access to high-quality machine translation systems and word alignment tools. We…

Computation and Language · Computer Science 2022-03-08 Tom Sherborne , Mirella Lapata

With the help of discrete neural audio codecs, large language models (LLM) have increasingly been recognized as a promising methodology for zero-shot Text-to-Speech (TTS) synthesis. However, sampling based decoding strategies bring…

Computation and Language · Computer Science 2024-06-13 Bing Han , Long Zhou , Shujie Liu , Sanyuan Chen , Lingwei Meng , Yanming Qian , Yanqing Liu , Sheng Zhao , Jinyu Li , Furu Wei

Robustness is critical in zero-shot singing voice conversion (SVC). This paper introduces two novel methods to strengthen the robustness of the kNN-VC framework for SVC. First, kNN-VC's core representation, WavLM, lacks harmonic emphasis,…

Sound · Computer Science 2025-04-09 Keren Shao , Ke Chen , Matthew Baas , Shlomo Dubnov

The widespread adoption of speech-based online services raises security and privacy concerns regarding the data that they use and share. If the data were compromised, attackers could exploit user speech to bypass speaker verification…

Sound · Computer Science 2022-09-13 Ruibin Yuan , Yuxuan Wu , Jacob Li , Jaxter Kim

This paper introduces PFlow-VC, a conditional flow matching voice conversion model that leverages fine-grained discrete pitch tokens and target speaker prompt information for expressive voice conversion (VC). Previous VC works primarily…

Non-parallel voice conversion (VC) is a technique for learning the mapping from source to target speech without relying on parallel data. This is an important task, but it has been challenging due to the disadvantages of the training…

Sound · Computer Science 2019-04-10 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Nobukatsu Hojo

Recent research in zero-shot speech synthesis has made significant progress in speaker similarity. However, current efforts focus on timbre generalization rather than prosody modeling, which results in limited naturalness and…

Sound · Computer Science 2024-06-12 Yuepeng Jiang , Tao Li , Fengyu Yang , Lei Xie , Meng Meng , Yujun Wang

Given a piece of speech and its transcript text, text-based speech editing aims to generate speech that can be seamlessly inserted into the given speech by editing the transcript. Existing methods adopt a two-stage approach: synthesize the…

Sound · Computer Science 2021-09-14 Chuanxin Tang , Chong Luo , Zhiyuan Zhao , Dacheng Yin , Yucheng Zhao , Wenjun Zeng
‹ Prev 1 4 5 6 7 8 10 Next ›