中文
相关论文

相关论文: StyleStream: Real-Time Zero-Shot Voice Style Conve…

200 篇论文

In real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ensure noise robustness but suppress prosody richness, while…

音频与语音处理 · 电气工程与系统科学 2025-08-11 Yuepeng Jiang , Ziqian Ning , Shuai Wang , Chengjia Wang , Mengxiao Bi , Pengcheng Zhu , Zhonghua Fu , Lei Xie

Expressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of speech emotion, it is challenging to disentangle the emotional…

音频与语音处理 · 电气工程与系统科学 2022-07-22 Zongyang Du , Berrak Sisman , Kun Zhou , Haizhou Li

Face-based Voice Conversion (FVC) is a novel task that leverages facial images to generate the target speaker's voice style. Previous work has two shortcomings: (1) suffering from obtaining facial embeddings that are well-aligned with the…

声音 · 计算机科学 2024-09-05 Yan Rong , Li Liu

Language model (LM) based audio generation frameworks, e.g., AudioLM, have recently achieved new state-of-the-art performance in zero-shot audio generation. In this paper, we explore the feasibility of LMs for zero-shot voice conversion. An…

音频与语音处理 · 电气工程与系统科学 2023-08-22 Zhichao Wang , Yuanzhe Chen , Lei Xie , Qiao Tian , Yuping Wang

Streaming speech-to-text translation (StreamST) is the task of automatically translating speech while incrementally receiving an audio stream. Unlike simultaneous ST (SimulST), which deals with pre-segmented speech, StreamST faces the…

声音 · 计算机科学 2024-06-11 Sara Papi , Marco Gaido , Matteo Negri , Luisa Bentivogli

This paper proposes a zero-shot text-to-speech (TTS) conditioned by a self-supervised speech-representation model acquired through self-supervised learning (SSL). Conventional methods with embedding vectors from x-vector or global style…

声音 · 计算机科学 2023-12-19 Kenichi Fujita , Takanori Ashihara , Hiroki Kanagawa , Takafumi Moriya , Yusuke Ijima

Zero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker's voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and modeling these tokens with the language model, recent language…

Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, producing both modalities simultaneously still typically relies on…

Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker traits persist-remains a challenge, especially in neural…

音频与语音处理 · 电气工程与系统科学 2025-07-15 Shivam Mehta , Yingru Liu , Zhenyu Tang , Kainan Peng , Vimal Manohar , Shun Zhang , Mike Seltzer , Qing He , Mingbo Ma

In this paper, we propose a model which can generate a singing voice from normal speech utterance by harnessing zero-shot, many-to-many style transfer learning. Our goal is to give anyone the opportunity to sing any song in a timely manner.…

音频与语音处理 · 电气工程与系统科学 2024-05-09 Amit Eliav , Aaron Taub , Renana Opochinsky , Sharon Gannot

Traditional studies on voice conversion (VC) have made progress with parallel training data and known speakers. Good voice conversion quality is obtained by exploring better alignment modules or expressive mapping functions. In this study,…

音频与语音处理 · 电气工程与系统科学 2022-04-01 Jiachen Lian , Chunlei Zhang , Dong Yu

Normally, a system that translates speech into text consists of separate modules for speech recognition and text-to-text translation. Combining those tasks into a SpeechLLM promises to exploit paralinguistic information in the speech and to…

计算与语言 · 计算机科学 2026-05-15 Titouan Parcollet , Shucong Zhang , Xianrui Zheng , Rogier C. van Dalen

Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide…

We propose a first streaming accent conversion (AC) model that transforms non-native speech into a native-like accent while preserving speaker identity, prosody and improving pronunciation. Our approach enables stream processing by…

计算与语言 · 计算机科学 2025-06-23 Tuan-Nam Nguyen , Ngoc-Quan Pham , Seymanur Akti , Alexander Waibel

While most research into speech synthesis has focused on synthesizing high-quality speech for in-dataset speakers, an equally essential yet unsolved problem is synthesizing speech for unseen speakers who are out-of-dataset with limited…

声音 · 计算机科学 2023-08-28 Wenbin Wang , Yang Song , Sanjay Jha

Zero-shot multi-speaker TTS aims to synthesize speech with the voice of a chosen target speaker without any fine-tuning. Prevailing methods, however, encounter limitations at adapting to new speakers of out-of-domain settings, primarily due…

声音 · 计算机科学 2024-03-06 Yejin Jeon , Yunsu Kim , Gary Geunbae Lee

In recent years, diffusion-based generative models have demonstrated remarkable performance in speech conversion, including Denoising Diffusion Probabilistic Models (DDPM) and others. However, the advantages of these models come at the cost…

声音 · 计算机科学 2025-06-03 Pengyu Ren , Wenhao Guan , Kaidi Wang , Peijie Chen , Qingyang Hong , Lin Li

YourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS. Our method builds upon the VITS model and adds several novel modifications for zero-shot multi-speaker and multilingual training. We achieved…

In the evolving domain of text-to-image generation, diffusion models have emerged as powerful tools in content creation. Despite their remarkable capability, existing models still face challenges in achieving controlled generation with a…

计算机视觉与模式识别 · 计算机科学 2024-02-22 Jaeseok Jeong , Junho Kim , Yunjey Choi , Gayoung Lee , Youngjung Uh

We introduce VoiceFilter-Lite, a single-channel source separation model that runs on the device to preserve only the speech signals from a target user, as part of a streaming speech recognition system. Delivering such a model presents…

音频与语音处理 · 电气工程与系统科学 2020-09-10 Quan Wang , Ignacio Lopez Moreno , Mert Saglam , Kevin Wilson , Alan Chiao , Renjie Liu , Yanzhang He , Wei Li , Jason Pelecanos , Marily Nika , Alexander Gruenstein